The request almost always arrives in the same shape: "we retype stacks of PDFs here every day — can't AI read those?" The answer is nearly always yes. Which is exactly why it is the wrong question.
A model can read them. The question that matters is: how often is it wrong, on which field, and who notices before it lands in your administration. That is not a model question but a process question — and you can measure it before you buy or build anything.
Start with the flow, not with the model
Automating a document flow starts by picking one flow: one type of document, one direction, one system where it ends up. Not "our administration", but "order confirmations from these twelve customers, into our order system".
At Boogaard Textiles, orders came in with customer data and print work as PDFs. The win was not in the reading itself — it was that the whole chain ended somewhere: cut lines generated, placed paper-efficiently into the print program, and a dashboard with barcode scanning that tracked where every order was. A day's work became twenty minutes. Had I only read the PDF and emailed the result, nothing would have been solved: someone would still be retyping it, with an extra step in between.
So start by counting. How many documents per week, how many different layouts, how many senders, and what happens to the result today. That counting takes an afternoon and decides whether the rest is worth it.
Measure per field, not per document
"95% accurate" is a number without meaning. Accurate on what? An IBAN that is wrong once in two hundred documents is unusable. A free-text reference that is wrong once in twenty may be perfectly fine — someone spots that immediately.
So build your own test set before you set anything up:
- Pull fifty to a hundred real documents from your own archive. Not the neat examples — specifically the skewed scans, the ones with a stamp over the amount, the sender who does everything differently.
- Type out by hand, once, what every field should have said. That is your ground truth, and it stays usable even if you later switch supplier or model.
- Score per field in three categories: correct, wrong, and not found.
That third category is the reason not to score per document. "Not found" is cheap: the system admits it does not know and hands the case to a person. "Wrong, but confident" is expensive, because it goes into your administration unseen.
The most expensive error is the silent one
An extraction that politely says it is unsure costs half a minute of attention. An extraction that writes away a wrong amount with the same ease costs a credit note, a phone call and some trust. So design the process around that:
- A threshold per field. Below it, the document goes to a review screen instead of the target system.
- The original alongside. Whoever reviews needs to see the PDF next to the filled-in field. Checking should take seconds, not minutes.
- Validate arithmetic instead of reading it. If a total has to follow from line items, add the lines up and compare. A sum that does not add up is a signal no model gives you.
- Log every correction. Which field, which sender, what it was and what it became. After a few weeks that log points precisely at what needs fixing.
That last one is a lesson from an entirely different project: a validation engine that found and categorised over 30,000 errors in an ERP environment (built as an employee at Agron). The value was not in the count — it was in the categorisation, because that turned "our data is messy" into a worklist you could start at the top of. Document extraction works the same way: without a correction log, all you know is that it sometimes goes wrong.
What actually breaks in practice
These never show up in a demo and always show up in production within a month:
- Skewed scans, phone photos, a stamp or initials across an amount.
- Handwriting. Do not count on it, however good the demo was.
- A regular sender changing their layout without telling anyone. You want an alert on that: "this looks different from last time".
- One PDF holding two invoices, or one invoice spread over five pages.
- The same document arriving twice. Duplicates are caught with a key (sender + number + amount), not with AI.
Not everything needs AI
The cheapest document automation is the document flow you avoid. If the sender can send a file instead of a PDF, or if you can reach their system, then an integration is more exact, faster and cheaper than any extraction. The broader question of what to tackle first is in process automation for SMEs.
Use AI where the source is genuinely unstructured and the sender is not going to change anything. That is often the case — but it is a choice, not a starting point.
How I approach it
One flow, a test set from your own archive, scores per field, a review screen for everything below the threshold, and a log of every correction. What I build is custom work around your existing systems — see AI & automation and software development — and the processing runs on infrastructure under our own management. If the documents contain personal data, a data processing agreement belongs with it; more on that in keeping your data in-house.
And if someone else puts a proposal on your table, ask these three questions: which documents was this measured on, per which field, and what happens to whatever falls below the threshold. A supplier without an answer has not measured it.