LLM and RAG engineering Guide
AI Document Processing in 2026: What It Reads Well, Where It Fails and How to Deploy It
Models now read a clean page almost perfectly. The trouble starts at the field level, where a wrong invoice number looks exactly like a right one. This guide sets out what independent tests and named deployments show, which documents will stop arriving as images, and how to build a pipeline that catches its own mistakes.
For CFOs, COOs, CTOs and heads of operations deciding how to automate invoices, claims, onboarding files, trade documents or medical records with AI.
The short answer
Use AI document processing for the documents that will keep arriving as scans and PDFs, and design it around checking rather than reading. Specialized models parse pages well, but field accuracy on real invoices runs from about 73% to 91%, so check every value in code, send uncertain fields to people, take e-invoices as data and keep a person on adverse decisions.
Key takeaways
- Small specialized parsing models lead the independent OmniDocBench at 96.91, ahead of Gemini 3 Pro at 92.91, so sending every page to a frontier model is not the most accurate choice.1
- Field extraction is harder than page reading: GPT-4o got 73.3% of fields right on a real-invoice benchmark, and the best key-field score on the IDP Leaderboard is 91.1%.23
- Language models fail plausibly: they rewrite short words up to 10% of the time and invent values on degraded scans, so errors look as clean as correct output.45
- A model's own confidence barely separates right from wrong answers (ROC AUC 0.69 to 0.74); engineered signals reached 99.1% accuracy on the 80% of fields they accepted.2
- Every named deployment that publishes results keeps people on the exceptions, and Allianz says its payout decisions are never automated.678
- E-invoicing mandates in Germany, Belgium, Poland and France turn most domestic invoices into structured data by 2028, so new OCR for those flows has a short life.9101112
Most document AI projects are bought on a demo: a messy invoice goes in, clean JSON comes out. The demo answers the easy question. Reading a page is close to solved, and the cost of reading it is a fraction of a cent. What decides whether a deployment pays is whether it can find the values it got wrong, how many documents still need a person, and whether the documents it was built for will keep arriving as images at all. This guide covers what independent benchmarks and named deployments show, how language models fail, what the work costs, which rules apply, and how to build a pipeline that checks itself.
- 73.3%of fields GPT-4o extracted correctly on a benchmark of real invoices2
- 32.6%of invoices processed with no human touch on average, against 49.2% at the best performers13
- 2028the year every German business must issue structured e-invoices, after which a plain PDF no longer counts9
How accurate AI document reading is
On OmniDocBench, an independent benchmark of 1,651 PDF pages across ten document types maintained by OpenDataLab, the leaders are small specialized models. TeleOCR scores 96.91 and PaddleOCR-VL-1.6 96.34, while Gemini 3 Pro scores 92.91, GPT-5.2 86.59 and Mistral OCR 85.66.1 Harder tests pull the ceiling down. On olmOCR-Bench, more than 7,000 pass-or-fail tests across 1,400 documents, the best system scores 83.1, and olmOCR itself passes 99.7% of basic tests but only 47.7% on old scans.14 The Allen Institute for AI built both the benchmark and olmOCR, a conflict that applies to most leaderboards in this field.
Business fields are where the numbers fall. On DocILE, a benchmark of real invoices with 55 field types, GPT-4o extracted 73.3% of fields correctly, a 26% failure rate.2 The best key-field score on the IDP Leaderboard is Gemini 3 Flash at 91.1%, measured by edit distance, which gives partial credit to an invoice number that is one digit off.3 Nanonets sponsors that leaderboard and ranks its own model first overall. On ExtractBench, 4,869 pages across 370 enterprise documents, commercial models did well on short documents but "often truncate record lists on long ones."16 A March 2026 comparison found classic OCR pipelines "generally more reliable for handwriting and for long or multi-page documents," and vision-language models stronger on multilingual text and rich layouts.17
Vendor figures shrink most when someone else measures. Mistral reported 94.89 for Mistral OCR on its own "text-only" test set of publication papers and web PDFs, with no sample size; independent benchmarks score the same model 72.0 and 85.66, and Mistral now marks it as no longer maintained.15141 A 99% claim should come with three answers: per character, per field or per document, on which documents, and how many.
How language models misread
Classic OCR fails noisily, with garbled characters a reviewer can spot. Language models fail fluently. A 2026 study of 15 systems on 1,455 documents found that "short words (4-6 characters) are rewritten up to 10% of the time," and that unusual words raised error rates by up to 6.9 points for general models against under 0.8 points for traditional OCR.4 Invoice numbers, policy IDs, dates and amounts are exactly that length. On degraded ID cards and invoices, models lean on "linguistic priors" and produce plausible values the image does not support; a 7-billion-parameter model trained to abstain gained 22 points of hallucination-free accuracy over GPT-4o.5
There is an old precedent. In 2013 a researcher found Xerox WorkCentre scanners changing digits in scanned building plans, with sixes turned into eights, because the compression algorithm swapped image patches it judged similar.18 The copies looked perfect. That is the failure to design for: output that is clean, confident and wrong.
Confidence scores do not tell you which fields are wrong
The obvious fix is to send low-confidence fields to a person. On real invoices it barely works. Using GPT-4o on DocILE, token log-probabilities separated right from wrong answers with a ROC AUC of 0.705, the model's own stated confidence 0.692 and agreement across five runs 0.744, where 0.5 is chance.2 The authors' explanation is plain: "A frontier LLM confidently transcribing OCR noise produces high log-probabilities for a wrong answer." A model that combined disagreement between runs with OCR, image-quality and layout signals reached 0.928, and by sending the least certain 20% of fields to people it raised accuracy on the rest to 99.1%.2 That is one paper from one vendor on one model, so treat it as a demonstration. The lesson holds: routing works when the signals are built and calibrated on your own documents.
The cheaper checks sit in code. Structured outputs guarantee shape and nothing else: OpenAI says its models "always" follow the supplied schema, and Anthropic's schema support drops numeric and length limits such as minimum and maximum.1920 So the rules live after the model: line items that sum to the subtotal, tax arithmetic, valid dates, IBAN and tax-ID checksums, supplier and customer master data, and matching the invoice to the purchase order and the goods receipt.21 Azure Content Understanding and Bedrock Data Automation both return confidence scores and grounding that ties each value to a region of the page, which lets a reviewer confirm a field in one glance.2223
What named deployments show
| Organization | Documents | What it reports | Where people stay |
|---|---|---|---|
| Uber | Supplier invoices, read by OCR then a language model | 90% overall accuracy; 35% of invoices near-perfect and 65% above 80%; handling time down 70%6 | Reviewers compare the PDF and the extracted data side by side |
| Rocket Close | About 2,000 title packages a day, 75 pages each | 89.71% accuracy across more than 44,000 fields; 30 minutes per package cut to under 27 | Not described |
| Allianz | Food spoilage claims after storms, seven AI agents | 80% less claim processing and settlement time8 | "Payout decisions are never automated" |
| HSBC | Trade finance documents in the UK, Hong Kong and the UAE | No figures published24 | Human judgment for "understanding exceptions and making the decisions that matter" |
| C.H. Robinson | Carrier and customer emails, orders and quotes | More than 3 million tasks and 1 million orders by AI; loads accepted in under 90 seconds against up to four hours25 | People take complex shipments |
| UK Department for Work and Pensions | About 25,000 scanned letters a day | Flags people who may need urgent help; no accuracy published26 | A daily list goes to trained staff; the tool makes no benefit decisions |
Uber's figures are the most candid. Overall accuracy is 90%, but only 35% of invoices came out near-perfect, and suppliers whose field accuracy falls below a threshold are prioritized for closer profiling and model work.6 Accuracy is a distribution across suppliers and layouts, and the tail is where the review effort goes. The UK Home Office found the same in its asylum pilots: a summarizing tool saved 23 minutes a case, with "occasional inaccuracies in summaries" and no source references.27
Programs also fail on intake. The IRS receives about 90 million paper documents a year, and paper returns were 6% of individual returns but 72% of processing costs in the 2025 season. By May 2025 contractors had scanned 5% of 9.8 million paper forms, and the next contractor about 7% of 5.7 million.28 Scanning capacity, procurement and document preparation decide outcomes as often as model accuracy does.
Where automated reading meets a decision
The serious failures so far are about decisions built on automated review. ProPublica reported that Cigna doctors denied more than 300,000 payment requests in two months using an automated method, spending an average of 1.2 seconds on each.29 A court let fiduciary claims proceed because the plan promised medical necessity review by a medical director.30 In the UnitedHealth case over its nH Predict tool, a federal court in 2026 ordered documents going back to January 2017.31 The pattern for buyers is practical: a person who signs in a second and a half is no defense, and the promise made to customers about human review is the standard a court applies.
| Rule | What it covers | What it means for document AI |
|---|---|---|
| EU AI Act, Article 6(3) and Annex III | A system doing "a narrow procedural task," such as one that "transforms unstructured data into structured data," is not high-risk32 | Extraction alone is usually outside the high-risk rules; feeding credit, insurance pricing or benefit decisions is not |
| EU Digital Omnibus on AI | Annex III high-risk duties apply from December 2, 202733 | About 14 months to add oversight, logging and risk management where extraction drives those decisions |
| GDPR Article 22 and the SCHUFA ruling | An automated score counts as a decision when a lender or other third party "draws strongly" on it34 | A recommendation a person always follows can still be an automated decision |
| California SB 1120 | Medical necessity "shall be made only by a licensed physician or a licensed health care professional"; AI may not supplant that decision35 | Extraction and summaries can assist a reviewer; they cannot deny care |
| CMS WISeR model | AI-assisted prior authorization in six states; a clinician reviews a denial before it is final36 | The same design: automate the approval path, keep people on denials |
| CMS-0057-F | Prior authorization decisions within 72 hours or 7 calendar days, with specific denial reasons, and FHIR APIs from 202737 | Requests move to structured data; attached clinical documents remain to be read |
| India DPDP Rules, 2025 | Notified November 2025 with an eighteen-month phased start38 | Consent, purpose and security duties cover every scanned form holding personal data |
| EU Anti-Money Laundering Regulation | Allows remote identification with electronic ID at "substantial" or "high" assurance39 | More EU onboarding will arrive with no document image at all |
Some documents will stop arriving as images
The biggest change to the business case is regulatory. In a growing list of countries a plain PDF stops being a valid invoice between businesses. Germany's finance ministry says a PDF sent by email no longer counts as an e-invoice; every business has had to receive structured invoices since January 1, 2025, and the transition for issuers ends after 2026, or after 2027 for those with turnover up to €800,000, with invoices up to €250 exempt.9
| Market | Structured e-invoicing between businesses | Effect on invoice reading |
|---|---|---|
| India | E-invoices above the turnover threshold; businesses at ₹10 crore or more must report within 30 days since April 1, 202540 | Large suppliers' invoices already arrive as registered data |
| Germany | Receipt mandatory since 2025; issuing from 2027 for larger firms and 2028 for all9 | Domestic invoices become XML by 2028; hybrid formats carry the data inside the PDF |
| Belgium | Issue and receive through Peppol from January 1, 2026; real-time reporting from January 202810 | Live |
| Poland | National system from February 1, 2026 for large firms and April 1, 2026 for others11 | Live |
| France | Receipt for all and issuing for large and mid-sized firms from September 1, 2026; smaller firms from September 1, 202712 | Live for the largest suppliers |
| United Kingdom | Mandatory e-invoicing for all VAT invoices from 202941 | Announced; specifications to follow |
| EU cross-border | Digital reporting for cross-border trade between businesses from July 1, 203042 | Adopted |
For a buyer whose suppliers sit mainly in those markets, a new OCR system for domestic invoices is a pipeline with a short life. The durable work is validating structured data, matching it to orders and receipts, and handling exceptions. Other documents move more slowly. Electronic bills of lading reached only about 5% of container trade in 2024, so trade documents will stay paper and scans for years, which is why HSBC is investing in reading them.4324 Claims evidence, clinical attachments, identity documents outside the EU's eID route and decades of archives will also stay images.
What it costs, and where the money goes
Reading is cheap. Google's Enterprise Document OCR lists $1.50 per 1,000 pages after the first 1,000, its Layout Parser $10 and its Custom Extractor $30 per 1,000 pages, and AWS's Textract example prices plain text detection at $0.0015 a page.4445 Units hide traps: Google's specialized parsers bill a count as "up to 10 pages," so a stream of one-page invoices costs ten times the per-page rate.44 Self-hosting goes lower still; the olmOCR team estimates a million pages for $176 on its own model, against over $6,240 for GPT-4o.46
Set that against the process. Ardent Partners puts the average cost of processing an invoice at $9.84, with an 18.4% exception rate and 8.2 days per invoice.47 The model is a rounding error on that figure. The business case turns on how many documents reach a person, how quickly each one is cleared, and how often a posted value turns out to be wrong.
Licenses and data handling narrow the shortlist before accuracy does. Marker's model weights are free only for organizations under $5 million in funding or revenue.48 MinerU requires a commercial license above 100 million monthly users or $20 million monthly revenue, and attribution for online services.49 The dots.ocr agreement bars extracting personal data protected by laws such as GDPR or HIPAA without consent and is governed by Chinese law.50 Among managed services, Azure Document Intelligence deletes inputs and results 24 hours after analysis, while Textract may use document inputs to improve AWS services unless the customer opts out.5152 OpenAI keeps abuse-monitoring logs for up to 30 days by default, and Anthropic's batch interface sits outside its zero-retention arrangement.5354
Managed document service
Prebuilt models in your cloud
- Fast start on common forms
- Confidence and page grounding included
- Fees per page, by feature
- Retention and training defaults to check
Frontier model and a prompt
Send the page, get JSON
- Any layout works on day one
- Errors look as clean as answers
- Confidence close to uninformative
- Long files truncated without warning
Checked pipeline
How we advise
- Structured inputs taken as data
- Specialized parser, then model extraction
- Every value checked in code and grounded
- Exceptions to a measured review queue
How to deploy AI document processing
- Sort the inbox by type and future format Count volumes by document type and source, and mark which will arrive as structured data under a mandate. Build reading for the streams that will stay images.
- Take structured inputs as data Accept e-invoices, Peppol messages and the XML inside hybrid PDFs directly, and validate them against the order and the supplier record.
- Label a sample before choosing a model Hand-label a few hundred documents per type, including bad scans and long files, and measure exact field accuracy for each candidate on them.
- Check every value in code Make totals add up, dates and IDs pass format and checksum tests, and every value match master data, the purchase order or a second document.
- Route by evidence Send a field to a person when it fails a check, disagrees between two readings or cannot be grounded on the page, and set thresholds from the labeled sample.
- Keep people on decisions, with time to decide Automate the approval path, never the adverse one, and track time per review and overturn rates so oversight stays real.
- Keep the record Store the source document, model version, prompt, output, checks and reviewer action for as long as tax and sector rules require the document itself.
Questions before signing a document AI contract
- What exact field accuracy does it reach on a labeled sample of our own documents, including bad scans?
- Which of our document streams will arrive as structured data within two years?
- How is each value checked, and can a reviewer see where on the page it came from?
- What share of documents will reach a person, and what does each one cost to clear?
- Who makes adverse decisions, and how long do they spend on each?
- Where are documents processed and stored, for how long, and are they used for training?
Document reading will keep getting better and cheaper, and the gains will show up first on the benchmarks that are easiest to run. The deployments that pay are built on the less visible parts: a labeled sample of real documents, checks that catch a clean-looking wrong value, structured data taken as data, and a person with the time to decide what matters.
This is how we approach AI development: measure on your own documents first, check every extracted value before it posts, and put people where the decisions are. On one dairy collection network, replacing paper registers with records captured at the point of entry cut payment calculation discrepancies by 92%.
Questions leaders ask
What is AI document processing?
AI document processing, also called intelligent document processing, uses OCR and language or vision models to classify documents such as invoices, claims, onboarding files and contracts, extract their fields into structured data, check the values and pass the result to business systems. Documents that fail a check go to a person for review.
How accurate is AI document extraction?
Page reading is highly accurate on clean documents, with the best specialized models scoring about 96 on OmniDocBench. Field extraction is lower: GPT-4o got 73.3% of fields right on real invoices, and the best key-field score on the IDP Leaderboard is 91.1%. Old scans, handwriting, long tables and multi-page files are the weak spots.
Can we trust a model's confidence score to route documents to people?
Not on its own. In a study of GPT-4o on real invoices, token probabilities, stated confidence and agreement across runs all separated right from wrong answers poorly. Routing works when it combines failed checks, disagreement between two readings and page grounding, with thresholds set on a labeled sample of your own documents.
Do we still need OCR if we use a large language model?
Usually yes. Specialized parsing models beat frontier models on independent page-parsing benchmarks, and classic OCR pipelines are more reliable on handwriting and long documents. A common design uses a parser or OCR layer for text and positions, then a language model to extract fields, with the OCR positions used to confirm each value on the page.
How much does AI document processing cost?
Reading is cheap: managed OCR lists around $1.50 per 1,000 pages, and field extraction services tens of dollars per 1,000. The larger cost is people handling exceptions; Ardent Partners puts the average cost of processing an invoice at $9.84. Plan the budget around the share of documents that reach a person.
Will e-invoicing make invoice OCR unnecessary?
For domestic invoices in markets with mandates, largely yes. Germany, Belgium, Poland and France move most business invoices to structured formats by 2028, and the UK from 2029. Invoices from suppliers elsewhere, small suppliers, receipts and expense claims will keep arriving as images, as will trade documents, claims evidence and medical records.
Is AI document processing high-risk under the EU AI Act?
Extraction on its own usually is not. The Act treats a system that transforms unstructured data into structured data as a narrow procedural task. It becomes high-risk when it drives an Annex III decision such as creditworthiness, life or health insurance pricing or eligibility for public benefits, with those duties applying from December 2, 2027.
Sources
- OmniDocBench: benchmarking diverse PDF document parsingOpenDataLab, 2026
- ExtractConf: calibrated confidence for LLM document extractionarXiv, June 2026
- IDP Leaderboard: task detailsIDP Leaderboard (sponsored by Nanonets)
- Do VLMs Read or Rewrite? Faithfulness in OCRarXiv, 2026
- KIE-HVQA: OCR hallucination in degraded documentsarXiv, NeurIPS 2025
- Advancing Invoice Document Processing at Uber using GenAIUber Engineering, April 2025
- Rocket Close transforms mortgage document processing with Amazon Bedrock and Amazon TextractAWS Machine Learning Blog, April 2026
- When the storm clears, so should the claim queueAllianz, November 3, 2025
- FAQ: Einführung der E-RechnungGerman Federal Ministry of Finance
- Belgium's 2026 e-invoicing regulations explainedVertex
- Poland announces new timeline for mandatory e-invoicingEY
- France confirms September 1, 2026 e-invoicing deadlineSovos
- AI in accounts payable: the metrics that matterTungsten Automation, April 15, 2025
- olmOCR and olmOCR-BenchAllen Institute for AI, 2025
- Mistral OCRMistral AI, March 6, 2025
- ExtractBench: schema-guided enterprise document extractionarXiv, July 2026
- DISCO: OCR pipelines and vision-language models on document understandingarXiv, March 2026
- Xerox copier flaw means dodgy numbers and dangerous designsThe Register, August 6, 2013
- Structured OutputsOpenAI API documentation
- Structured outputsClaude Developer Platform documentation
- Three-way matching policiesMicrosoft Learn, Dynamics 365 Finance
- What is Azure Content Understanding?Microsoft Learn
- Bedrock Data AutomationAWS documentation
- HSBC goes proprietary on AI trade doc checkingGlobal Trade Review, 2026
- AI performs over three million shipping tasksC.H. Robinson, April 16, 2025
- Whitemail insights and vulnerability scanner: algorithmic transparency recordGOV.UK, 2025
- Home Office to expand AI use in asylum decision-makingElectronic Immigration Network, 2025
- Paper return processing: digitization resultsTIGTA, February 2026
- How Cigna saves millions by having its doctors reject claims without reading themProPublica, March 25, 2023
- Court partially grants, partially denies Cigna's motion to dismiss AI claims review caseDigital Healthcare Law, May 14, 2025
- Federal court orders broad discovery against UHC in AI coverage denial lawsuitArentFox Schiff, March 2026
- Regulation (EU) 2024/1689, the AI ActEUR-Lex
- Regulation (EU) 2026/1744, Digital Omnibus on AIEUR-Lex, July 2026
- CJEU rules that a credit score constitutes automated decision-making under the GDPRA&O Shearman, 2023
- SB-1120 Health care coverage: utilization reviewCalifornia Legislative Information, 2024
- Is AI WISeR? CMS models AI-based prior authorization in six statesDinsmore, 2025
- CMS-0057-F: Interoperability and Prior Authorization final ruleFirely
- Digital Personal Data Protection Rules, 2025: phased commencementPress Information Bureau, November 2025
- Regulation (EU) 2024/1624 on preventing the use of the financial system for money launderingEUR-Lex
- Revised time limit for e-invoice reporting for businesses with AATO of ₹10 crores and aboveGSTN
- Promoting electronic invoicing across UK businesses and the public sector: consultation responseGOV.UK, November 26, 2025
- VAT in the Digital Age (ViDA)European Commission
- eBL adoption doubles to 5% but barriers to digitisation remain, DCSA findsGlobal Trade Review, 2025
- Document AI pricingGoogle Cloud
- Amazon Textract pricingAWS
- olmOCR: unlocking trillions of tokens in PDFs with vision language modelsarXiv, 2025
- State of ePayables: AP benchmarks and best-in-class performanceArdent Partners, January 2026
- Marker: commercial usageDatalab on GitHub
- MinerU licenseOpenDataLab on GitHub
- dots.ocr license agreementrednote-hilab on Hugging Face
- Data, privacy and security for Document IntelligenceMicrosoft Learn
- Amazon Textract FAQsAWS
- Data controls in the OpenAI platformOpenAI API documentation
- API and data retentionClaude Developer Platform documentation
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 5, 2026. No client data appears in our insights.