Document Intelligence
Every vendor demonstrates on a clean PDF. Your inbox contains a photograph of a form, taken at an angle, with a correction in biro.
- Typical duration
- Three to eight months
- Team shape
- Two to four senior engineers
- Starts with
- Your real document set

The awkward documents are the project
Extraction accuracy on well-structured digital documents has been a solved problem for a while. What decides whether a document project succeeds is the rest of the pile: scans, phone photographs, forms completed by hand, layouts a supplier changed without telling anyone, and the one document type that represents four per cent of volume and thirty per cent of the exceptions.
So the evaluation runs on a sample drawn from what you actually receive rather than on examples anyone selected, and it reports per document type rather than as one number. The build that follows is then mostly about the confidence threshold and the review queue — deciding which extractions go straight through and which a person checks, because that boundary is where the value and the risk both sit.
What usually goes wrong, and what we do instead
The pattern is consistent enough that we now insist on a real sample before quoting anything.
Where document projects fail
- Evaluated on clean examples, then met with photographs and scans
- One headline accuracy figure hiding a document type that fails badly
- Everything routed to review, so no work is actually saved
- Nothing routed to review, so errors reach the ledger unnoticed
- A supplier changes a layout and accuracy drops with no alert
How we build it instead
- A sample drawn from real intake, including the awkward tail
- Accuracy reported per document type and per field, never as one number
- A confidence threshold tuned to the cost of an error, with you
- A review queue designed as a workflow rather than a fallback
- Per-source accuracy monitoring, so a layout change raises an alert
What that involves
The work that decides whether extraction saves effort or moves it somewhere else.
Evaluation on your real intake
Not a curated sample. The pile as it arrives.
We score extraction on documents drawn from what you actually receive, and report per document type and per field. The awkward tail is where these projects fail, so it is where the evaluation looks first.
Scans and photographs
Skewed, shadowed and low-resolution capture, which is the majority of real intake in most processes.
Handwriting and forms
Fields completed by hand, corrected, or filled in outside the box they belong in.
Confidence thresholds
Where straight-through processing ends and review begins, tuned to what an error costs you.
Review queue design
A workflow a person can move through quickly, with the original beside the extraction.
Downstream integration
Extracted data written into the system that uses it, with the source document still reachable.
Per-source monitoring
Accuracy tracked by sender and layout, so a supplier changing a form raises an alert.
What you end up with
Stated from your side rather than ours.
An answer before the investment
The evaluation tells you what proportion of your real intake can go straight through, per document type, before you commit.
A threshold you chose
Where automation ends and review begins is your decision, priced against what an error actually costs rather than set by a default.
A review queue people can work
The exceptions arrive as a workflow with the original alongside, rather than as a rejected pile somebody has to re-key.
An audit trail back to the page
Every extracted value points at the document and the region it came from, which is what an auditor asks for.
How the pieces fit together
Extraction is one step in six, and the two after it decide whether the process actually gets faster.

What each part does
- Intake: documents arriving from every channel you actually use, stored unmodified so the original is always reachable.
- Classification: what kind of document this is, because extraction rules and accuracy differ sharply by type.
- Extraction: fields read from the page, each with a confidence score and a reference to the region it came from.
- The confidence gate: the threshold you set, deciding what goes straight through and what a person checks.
- Review queue: the exceptions, presented beside the original page so a correction takes seconds rather than minutes.
- Monitoring: accuracy tracked per sender and layout, so a supplier changing a form is an alert rather than a slow decline.
How the work runs
The same five stages as every engagement, applied to a process fed by paper.
Diagnose
What arrives, in what proportions, and what it costs you today to get it into a system.
Decide
The evaluation is run on real intake and the threshold agreed. This is where a no costs you weeks rather than a year.
Prove
One document type end to end in production, including its review queue and its downstream write.
Deliver
Document types added in slices, each evaluated before the next is scoped.
Hand over
Your team runs the evaluation and adjusts the thresholds with us watching, then without us.
What we build it with
Chosen for what has to be explained afterwards, and for who has to run it once we have gone.
Models
- Anthropic
- OpenAI
- Gemini
- PyTorch
Document processing
- Python
- Tesseract
- pgvector
- Elasticsearch
Services
- TypeScript
- Node.js
- PostgreSQL
- Temporal
Platform
- AWS
- Azure
- Docker
- Datadog
What you get, and when
Handed over as it is produced, not assembled at the end.
- Accuracy per document type and per field, on your real intake
- The straight-through proportion at several thresholds
- A written recommendation, which is allowed to be 'not yet'
- Source in repositories you own
- A pipeline you can re-run over a historical batch
- The review queue, as a workflow rather than an inbox
- Every extracted value traceable to a page and a region
- The threshold decision, recorded with the cost basis behind it
- Per-sender accuracy dashboards with alerting on decline
The rest of our AI work
Most AI engagements need one of these underneath them, and some need two.
Related work
Engagements where the awkward inputs were the whole problem.
Questions about document intelligence
The ones we are asked most often before a first conversation.
How accurate is the extraction?
Can it read handwriting?
Will we still need people checking documents?
Who decides what goes through automatically?
What happens when a supplier changes their form?
Our documents cannot leave our environment.
Drowning in documents?
Send us a sample of what actually arrives. We will tell you what proportion could go straight through.

