Skip to content
CipherCruCipherCru

Menu

AI Development

Document Intelligence

Every vendor demonstrates on a clean PDF. Your inbox contains a photograph of a form, taken at an angle, with a correction in biro.

Typical duration
Three to eight months
Team shape
Two to four senior engineers
Starts with
Your real document set
Extracted document fields under review beside the original scans

The awkward documents are the project

Extraction accuracy on well-structured digital documents has been a solved problem for a while. What decides whether a document project succeeds is the rest of the pile: scans, phone photographs, forms completed by hand, layouts a supplier changed without telling anyone, and the one document type that represents four per cent of volume and thirty per cent of the exceptions.

So the evaluation runs on a sample drawn from what you actually receive rather than on examples anyone selected, and it reports per document type rather than as one number. The build that follows is then mostly about the confidence threshold and the review queue — deciding which extractions go straight through and which a person checks, because that boundary is where the value and the risk both sit.

What usually goes wrong, and what we do instead

The pattern is consistent enough that we now insist on a real sample before quoting anything.

Where document projects fail

  • Evaluated on clean examples, then met with photographs and scans
  • One headline accuracy figure hiding a document type that fails badly
  • Everything routed to review, so no work is actually saved
  • Nothing routed to review, so errors reach the ledger unnoticed
  • A supplier changes a layout and accuracy drops with no alert

How we build it instead

  • A sample drawn from real intake, including the awkward tail
  • Accuracy reported per document type and per field, never as one number
  • A confidence threshold tuned to the cost of an error, with you
  • A review queue designed as a workflow rather than a fallback
  • Per-source accuracy monitoring, so a layout change raises an alert

What that involves

The work that decides whether extraction saves effort or moves it somewhere else.

Evaluation on your real intake

Not a curated sample. The pile as it arrives.

We score extraction on documents drawn from what you actually receive, and report per document type and per field. The awkward tail is where these projects fail, so it is where the evaluation looks first.

Scans and photographs

Skewed, shadowed and low-resolution capture, which is the majority of real intake in most processes.

Handwriting and forms

Fields completed by hand, corrected, or filled in outside the box they belong in.

Confidence thresholds

Where straight-through processing ends and review begins, tuned to what an error costs you.

Review queue design

A workflow a person can move through quickly, with the original beside the extraction.

Downstream integration

Extracted data written into the system that uses it, with the source document still reachable.

Per-source monitoring

Accuracy tracked by sender and layout, so a supplier changing a form raises an alert.

What you end up with

Stated from your side rather than ours.

  • An answer before the investment

    The evaluation tells you what proportion of your real intake can go straight through, per document type, before you commit.

  • A threshold you chose

    Where automation ends and review begins is your decision, priced against what an error actually costs rather than set by a default.

  • A review queue people can work

    The exceptions arrive as a workflow with the original alongside, rather than as a rejected pile somebody has to re-key.

  • An audit trail back to the page

    Every extracted value points at the document and the region it came from, which is what an auditor asks for.

How the pieces fit together

Extraction is one step in six, and the two after it decide whether the process actually gets faster.

A document pipeline from intake through classification and extraction to a confidence gate, a review queue and a downstream system
Stand-in artwork. The numbered legend describes the system rather than this picture.

What each part does

  1. Intake: documents arriving from every channel you actually use, stored unmodified so the original is always reachable.
  2. Classification: what kind of document this is, because extraction rules and accuracy differ sharply by type.
  3. Extraction: fields read from the page, each with a confidence score and a reference to the region it came from.
  4. The confidence gate: the threshold you set, deciding what goes straight through and what a person checks.
  5. Review queue: the exceptions, presented beside the original page so a correction takes seconds rather than minutes.
  6. Monitoring: accuracy tracked per sender and layout, so a supplier changing a form is an alert rather than a slow decline.

How the work runs

The same five stages as every engagement, applied to a process fed by paper.

Diagnose

What arrives, in what proportions, and what it costs you today to get it into a system.

Decide

The evaluation is run on real intake and the threshold agreed. This is where a no costs you weeks rather than a year.

Prove

One document type end to end in production, including its review queue and its downstream write.

Deliver

Document types added in slices, each evaluated before the next is scoped.

Hand over

Your team runs the evaluation and adjusts the thresholds with us watching, then without us.

What we build it with

Chosen for what has to be explained afterwards, and for who has to run it once we have gone.

Models

  • Anthropic
  • OpenAI
  • Gemini
  • PyTorch

Document processing

  • Python
  • Tesseract
  • pgvector
  • Elasticsearch

Services

  • TypeScript
  • Node.js
  • PostgreSQL
  • Temporal

Platform

  • AWS
  • Azure
  • Docker
  • Datadog

What you get, and when

Handed over as it is produced, not assembled at the end.

  • Accuracy per document type and per field, on your real intake
  • The straight-through proportion at several thresholds
  • A written recommendation, which is allowed to be 'not yet'
  • Source in repositories you own
  • A pipeline you can re-run over a historical batch
  • The review queue, as a workflow rather than an inbox
  • Every extracted value traceable to a page and a region
  • The threshold decision, recorded with the cost basis behind it
  • Per-sender accuracy dashboards with alerting on decline

Questions about document intelligence

The ones we are asked most often before a first conversation.

How accurate is the extraction?
It depends entirely on your documents, which is why we will not quote a figure here and are wary of anyone who does. The evaluation gives you accuracy per document type and per field on a sample of your real intake — and reporting it that way matters, because a headline number of ninety-something routinely hides one document type that fails almost completely.
Can it read handwriting?
Often, and considerably better than a few years ago, but this is exactly where a general claim is worthless. Block capitals in a bounded field are close to reliable; cursive in a free-text box, or a correction written over a printed value, is not. The evaluation tells you which of your fields fall on which side, and the answer usually shapes the threshold more than it shapes the model choice.
Will we still need people checking documents?
Yes, and a system claiming otherwise is one you should not deploy. The aim is not to remove review but to shrink it to the documents that genuinely need it, and to make each one fast — the extraction beside the original page, with the uncertain field highlighted. Where the saving comes from is the eighty per cent that never reaches a person at all.
Who decides what goes through automatically?
You do, and it is a commercial decision rather than a technical one. The evaluation gives you the straight-through rate at several thresholds and the error rate at each, and you choose based on what an error costs in your process — which for an invoice line and for a medical record are not remotely the same number. It can also differ per field, and often should.
What happens when a supplier changes their form?
Accuracy is monitored per sender and per layout precisely so this is an alert rather than a slow drift nobody notices. Modern extraction handles moderate layout change far better than the template-based systems it replaced, so a moved field is usually absorbed. A wholesale redesign will need attention, and you will know within a day rather than at a month-end reconciliation.
Our documents cannot leave our environment.
Then they will not. Open models running in your own infrastructure are a legitimate option for extraction, with some cost in accuracy on the hardest documents — which the evaluation will quantify for you rather than leave as a worry. Establishing that constraint is one of the first things we do, because it changes the architecture rather than being a setting.

Drowning in documents?

Send us a sample of what actually arrives. We will tell you what proportion could go straight through.

Strictly necessaryEssential for the site to function: page navigation, security, session management, and remembering the cookie choices you make here.
Always on
FunctionalRemembers choices you make, such as language, region or display preferences, so the site opens the way you left it.
Performance and analyticsPerformance and analytics cookies show us how the Website is used: which pages are visited, how long is spent on them, where visitors came from, and what errors occur. They are set by Google Analytics and by HubSpot, whose cookies also link the pages you viewed to any enquiry you later send us.
Marketing and targetingTracks browsing activity to measure advertising and show relevant ads. We set none of these today, and will not without your opt-in.