LLM and RAG engineering Guide

AI Document Processing in 2026: What It Reads Well, Where It Fails and How to Deploy It

Models now read a clean page almost perfectly. The trouble starts at the field level, where a wrong invoice number looks exactly like a right one. This guide sets out what independent tests and named deployments show, which documents will stop arriving as images, and how to build a pipeline that catches its own mistakes.

For CFOs, COOs, CTOs and heads of operations deciding how to automate invoices, claims, onboarding files, trade documents or medical records with AI.

Published
Reviewed
Reading time
15 min

The short answer

Use AI document processing for the documents that will keep arriving as scans and PDFs, and design it around checking rather than reading. Specialized models parse pages well, but field accuracy on real invoices runs from about 73% to 91%, so check every value in code, send uncertain fields to people, take e-invoices as data and keep a person on adverse decisions.

Key takeaways

  • Small specialized parsing models lead the independent OmniDocBench at 96.91, ahead of Gemini 3 Pro at 92.91, so sending every page to a frontier model is not the most accurate choice.1
  • Field extraction is harder than page reading: GPT-4o got 73.3% of fields right on a real-invoice benchmark, and the best key-field score on the IDP Leaderboard is 91.1%.23
  • Language models fail plausibly: they rewrite short words up to 10% of the time and invent values on degraded scans, so errors look as clean as correct output.45
  • A model's own confidence barely separates right from wrong answers (ROC AUC 0.69 to 0.74); engineered signals reached 99.1% accuracy on the 80% of fields they accepted.2
  • Every named deployment that publishes results keeps people on the exceptions, and Allianz says its payout decisions are never automated.678
  • E-invoicing mandates in Germany, Belgium, Poland and France turn most domestic invoices into structured data by 2028, so new OCR for those flows has a short life.9101112

Most document AI projects are bought on a demo: a messy invoice goes in, clean JSON comes out. The demo answers the easy question. Reading a page is close to solved, and the cost of reading it is a fraction of a cent. What decides whether a deployment pays is whether it can find the values it got wrong, how many documents still need a person, and whether the documents it was built for will keep arriving as images at all. This guide covers what independent benchmarks and named deployments show, how language models fail, what the work costs, which rules apply, and how to build a pipeline that checks itself.

  • 73.3%of fields GPT-4o extracted correctly on a benchmark of real invoices2
  • 32.6%of invoices processed with no human touch on average, against 49.2% at the best performers13
  • 2028the year every German business must issue structured e-invoices, after which a plain PDF no longer counts9

How accurate AI document reading is

On OmniDocBench, an independent benchmark of 1,651 PDF pages across ten document types maintained by OpenDataLab, the leaders are small specialized models. TeleOCR scores 96.91 and PaddleOCR-VL-1.6 96.34, while Gemini 3 Pro scores 92.91, GPT-5.2 86.59 and Mistral OCR 85.66.1 Harder tests pull the ceiling down. On olmOCR-Bench, more than 7,000 pass-or-fail tests across 1,400 documents, the best system scores 83.1, and olmOCR itself passes 99.7% of basic tests but only 47.7% on old scans.14 The Allen Institute for AI built both the benchmark and olmOCR, a conflict that applies to most leaderboards in this field.

Exhibit 1What the benchmarks measure, from page reading down to exact fields
  • Best specialized parser, page parsing (OmniDocBench)96.9
  • Gemini 3 Pro, page parsing (OmniDocBench)92.9
  • Best key-field score, partial credit (IDP Leaderboard)91.1%
  • Best system, pass-or-fail tests (olmOCR-Bench)83.1
  • GPT-4o, invoice fields correct (DocILE)73.3%
  • Mistral OCR on olmOCR-Bench, after claiming 94.89 on its own test set72.0
  • olmOCR on old scans (olmOCR-Bench)47.7%
Scores come from different benchmarks and metrics, so they are not like for like. The point is the direction: the closer a test gets to exact business fields and real scans, the lower the number. Sources: [1], [3], [14], [2], [15]

Business fields are where the numbers fall. On DocILE, a benchmark of real invoices with 55 field types, GPT-4o extracted 73.3% of fields correctly, a 26% failure rate.2 The best key-field score on the IDP Leaderboard is Gemini 3 Flash at 91.1%, measured by edit distance, which gives partial credit to an invoice number that is one digit off.3 Nanonets sponsors that leaderboard and ranks its own model first overall. On ExtractBench, 4,869 pages across 370 enterprise documents, commercial models did well on short documents but "often truncate record lists on long ones."16 A March 2026 comparison found classic OCR pipelines "generally more reliable for handwriting and for long or multi-page documents," and vision-language models stronger on multilingual text and rich layouts.17

Vendor figures shrink most when someone else measures. Mistral reported 94.89 for Mistral OCR on its own "text-only" test set of publication papers and web PDFs, with no sample size; independent benchmarks score the same model 72.0 and 85.66, and Mistral now marks it as no longer maintained.15141 A 99% claim should come with three answers: per character, per field or per document, on which documents, and how many.

How language models misread

Classic OCR fails noisily, with garbled characters a reviewer can spot. Language models fail fluently. A 2026 study of 15 systems on 1,455 documents found that "short words (4-6 characters) are rewritten up to 10% of the time," and that unusual words raised error rates by up to 6.9 points for general models against under 0.8 points for traditional OCR.4 Invoice numbers, policy IDs, dates and amounts are exactly that length. On degraded ID cards and invoices, models lean on "linguistic priors" and produce plausible values the image does not support; a 7-billion-parameter model trained to abstain gained 22 points of hallucination-free accuracy over GPT-4o.5

There is an old precedent. In 2013 a researcher found Xerox WorkCentre scanners changing digits in scanned building plans, with sixes turned into eights, because the compression algorithm swapped image patches it judged similar.18 The copies looked perfect. That is the failure to design for: output that is clean, confident and wrong.

Confidence scores do not tell you which fields are wrong

The obvious fix is to send low-confidence fields to a person. On real invoices it barely works. Using GPT-4o on DocILE, token log-probabilities separated right from wrong answers with a ROC AUC of 0.705, the model's own stated confidence 0.692 and agreement across five runs 0.744, where 0.5 is chance.2 The authors' explanation is plain: "A frontier LLM confidently transcribing OCR noise produces high log-probabilities for a wrong answer." A model that combined disagreement between runs with OCR, image-quality and layout signals reached 0.928, and by sending the least certain 20% of fields to people it raised accuracy on the rest to 99.1%.2 That is one paper from one vendor on one model, so treat it as a demonstration. The lesson holds: routing works when the signals are built and calibrated on your own documents.

The cheaper checks sit in code. Structured outputs guarantee shape and nothing else: OpenAI says its models "always" follow the supplied schema, and Anthropic's schema support drops numeric and length limits such as minimum and maximum.1920 So the rules live after the model: line items that sum to the subtotal, tax arithmetic, valid dates, IBAN and tax-ID checksums, supplier and customer master data, and matching the invoice to the purchase order and the goods receipt.21 Azure Content Understanding and Bedrock Data Automation both return confidence scores and grounding that ties each value to a region of the page, which lets a reviewer confirm a field in one glance.2223

What named deployments show

OrganizationDocumentsWhat it reportsWhere people stay
UberSupplier invoices, read by OCR then a language model90% overall accuracy; 35% of invoices near-perfect and 65% above 80%; handling time down 70%6Reviewers compare the PDF and the extracted data side by side
Rocket CloseAbout 2,000 title packages a day, 75 pages each89.71% accuracy across more than 44,000 fields; 30 minutes per package cut to under 27Not described
AllianzFood spoilage claims after storms, seven AI agents80% less claim processing and settlement time8"Payout decisions are never automated"
HSBCTrade finance documents in the UK, Hong Kong and the UAENo figures published24Human judgment for "understanding exceptions and making the decisions that matter"
C.H. RobinsonCarrier and customer emails, orders and quotesMore than 3 million tasks and 1 million orders by AI; loads accepted in under 90 seconds against up to four hours25People take complex shipments
UK Department for Work and PensionsAbout 25,000 scanned letters a dayFlags people who may need urgent help; no accuracy published26A daily list goes to trained staff; the tool makes no benefit decisions
Where accuracy is published it sits near 90%, and every system routes exceptions to people by design.

Uber's figures are the most candid. Overall accuracy is 90%, but only 35% of invoices came out near-perfect, and suppliers whose field accuracy falls below a threshold are prioritized for closer profiling and model work.6 Accuracy is a distribution across suppliers and layouts, and the tail is where the review effort goes. The UK Home Office found the same in its asylum pilots: a summarizing tool saved 23 minutes a case, with "occasional inaccuracies in summaries" and no source references.27

Programs also fail on intake. The IRS receives about 90 million paper documents a year, and paper returns were 6% of individual returns but 72% of processing costs in the 2025 season. By May 2025 contractors had scanned 5% of 9.8 million paper forms, and the next contractor about 7% of 5.7 million.28 Scanning capacity, procurement and document preparation decide outcomes as often as model accuracy does.

Where automated reading meets a decision

The serious failures so far are about decisions built on automated review. ProPublica reported that Cigna doctors denied more than 300,000 payment requests in two months using an automated method, spending an average of 1.2 seconds on each.29 A court let fiduciary claims proceed because the plan promised medical necessity review by a medical director.30 In the UnitedHealth case over its nH Predict tool, a federal court in 2026 ordered documents going back to January 2017.31 The pattern for buyers is practical: a person who signs in a second and a half is no defense, and the promise made to customers about human review is the standard a court applies.

RuleWhat it coversWhat it means for document AI
EU AI Act, Article 6(3) and Annex IIIA system doing "a narrow procedural task," such as one that "transforms unstructured data into structured data," is not high-risk32Extraction alone is usually outside the high-risk rules; feeding credit, insurance pricing or benefit decisions is not
EU Digital Omnibus on AIAnnex III high-risk duties apply from December 2, 202733About 14 months to add oversight, logging and risk management where extraction drives those decisions
GDPR Article 22 and the SCHUFA rulingAn automated score counts as a decision when a lender or other third party "draws strongly" on it34A recommendation a person always follows can still be an automated decision
California SB 1120Medical necessity "shall be made only by a licensed physician or a licensed health care professional"; AI may not supplant that decision35Extraction and summaries can assist a reviewer; they cannot deny care
CMS WISeR modelAI-assisted prior authorization in six states; a clinician reviews a denial before it is final36The same design: automate the approval path, keep people on denials
CMS-0057-FPrior authorization decisions within 72 hours or 7 calendar days, with specific denial reasons, and FHIR APIs from 202737Requests move to structured data; attached clinical documents remain to be read
India DPDP Rules, 2025Notified November 2025 with an eighteen-month phased start38Consent, purpose and security duties cover every scanned form holding personal data
EU Anti-Money Laundering RegulationAllows remote identification with electronic ID at "substantial" or "high" assurance39More EU onboarding will arrive with no document image at all
Status as of October 2026. The rules mostly leave the act of reading alone and govern what is decided with the result.

Some documents will stop arriving as images

The biggest change to the business case is regulatory. In a growing list of countries a plain PDF stops being a valid invoice between businesses. Germany's finance ministry says a PDF sent by email no longer counts as an e-invoice; every business has had to receive structured invoices since January 1, 2025, and the transition for issuers ends after 2026, or after 2027 for those with turnover up to €800,000, with invoices up to €250 exempt.9

MarketStructured e-invoicing between businessesEffect on invoice reading
IndiaE-invoices above the turnover threshold; businesses at ₹10 crore or more must report within 30 days since April 1, 202540Large suppliers' invoices already arrive as registered data
GermanyReceipt mandatory since 2025; issuing from 2027 for larger firms and 2028 for all9Domestic invoices become XML by 2028; hybrid formats carry the data inside the PDF
BelgiumIssue and receive through Peppol from January 1, 2026; real-time reporting from January 202810Live
PolandNational system from February 1, 2026 for large firms and April 1, 2026 for others11Live
FranceReceipt for all and issuing for large and mid-sized firms from September 1, 2026; smaller firms from September 1, 202712Live for the largest suppliers
United KingdomMandatory e-invoicing for all VAT invoices from 202941Announced; specifications to follow
EU cross-borderDigital reporting for cross-border trade between businesses from July 1, 203042Adopted
Invoices from suppliers in markets without a mandate, small suppliers below thresholds, receipts and expense claims will keep arriving as images.

For a buyer whose suppliers sit mainly in those markets, a new OCR system for domestic invoices is a pipeline with a short life. The durable work is validating structured data, matching it to orders and receipts, and handling exceptions. Other documents move more slowly. Electronic bills of lading reached only about 5% of container trade in 2024, so trade documents will stay paper and scans for years, which is why HSBC is investing in reading them.4324 Claims evidence, clinical attachments, identity documents outside the EU's eID route and decades of archives will also stay images.

What it costs, and where the money goes

Reading is cheap. Google's Enterprise Document OCR lists $1.50 per 1,000 pages after the first 1,000, its Layout Parser $10 and its Custom Extractor $30 per 1,000 pages, and AWS's Textract example prices plain text detection at $0.0015 a page.4445 Units hide traps: Google's specialized parsers bill a count as "up to 10 pages," so a stream of one-page invoices costs ten times the per-page rate.44 Self-hosting goes lower still; the olmOCR team estimates a million pages for $176 on its own model, against over $6,240 for GPT-4o.46

Set that against the process. Ardent Partners puts the average cost of processing an invoice at $9.84, with an 18.4% exception rate and 8.2 days per invoice.47 The model is a rounding error on that figure. The business case turns on how many documents reach a person, how quickly each one is cleared, and how often a posted value turns out to be wrong.

Licenses and data handling narrow the shortlist before accuracy does. Marker's model weights are free only for organizations under $5 million in funding or revenue.48 MinerU requires a commercial license above 100 million monthly users or $20 million monthly revenue, and attribution for online services.49 The dots.ocr agreement bars extracting personal data protected by laws such as GDPR or HIPAA without consent and is governed by Chinese law.50 Among managed services, Azure Document Intelligence deletes inputs and results 24 hours after analysis, while Textract may use document inputs to improve AWS services unless the customer opts out.5152 OpenAI keeps abuse-monitoring logs for up to 30 days by default, and Anthropic's batch interface sits outside its zero-retention arrangement.5354

Exhibit 2Three ways to build it

Managed document service

Prebuilt models in your cloud

  • Fast start on common forms
  • Confidence and page grounding included
  • Fees per page, by feature
  • Retention and training defaults to check

Frontier model and a prompt

Send the page, get JSON

  • Any layout works on day one
  • Errors look as clean as answers
  • Confidence close to uninformative
  • Long files truncated without warning

Checked pipeline

How we advise

  • Structured inputs taken as data
  • Specialized parser, then model extraction
  • Every value checked in code and grounded
  • Exceptions to a measured review queue
Most production systems combine the first and third: a managed or open parser for reading, with checks, routing and review built around it.

How to deploy AI document processing

  1. Sort the inbox by type and future format Count volumes by document type and source, and mark which will arrive as structured data under a mandate. Build reading for the streams that will stay images.
  2. Take structured inputs as data Accept e-invoices, Peppol messages and the XML inside hybrid PDFs directly, and validate them against the order and the supplier record.
  3. Label a sample before choosing a model Hand-label a few hundred documents per type, including bad scans and long files, and measure exact field accuracy for each candidate on them.
  4. Check every value in code Make totals add up, dates and IDs pass format and checksum tests, and every value match master data, the purchase order or a second document.
  5. Route by evidence Send a field to a person when it fails a check, disagrees between two readings or cannot be grounded on the page, and set thresholds from the labeled sample.
  6. Keep people on decisions, with time to decide Automate the approval path, never the adverse one, and track time per review and overturn rates so oversight stays real.
  7. Keep the record Store the source document, model version, prompt, output, checks and reviewer action for as long as tax and sector rules require the document itself.

Questions before signing a document AI contract

  • What exact field accuracy does it reach on a labeled sample of our own documents, including bad scans?
  • Which of our document streams will arrive as structured data within two years?
  • How is each value checked, and can a reviewer see where on the page it came from?
  • What share of documents will reach a person, and what does each one cost to clear?
  • Who makes adverse decisions, and how long do they spend on each?
  • Where are documents processed and stored, for how long, and are they used for training?

Document reading will keep getting better and cheaper, and the gains will show up first on the benchmarks that are easiest to run. The deployments that pay are built on the less visible parts: a labeled sample of real documents, checks that catch a clean-looking wrong value, structured data taken as data, and a person with the time to decide what matters.

This is how we approach AI development: measure on your own documents first, check every extracted value before it posts, and put people where the decisions are. On one dairy collection network, replacing paper registers with records captured at the point of entry cut payment calculation discrepancies by 92%.

Questions leaders ask

What is AI document processing?

AI document processing, also called intelligent document processing, uses OCR and language or vision models to classify documents such as invoices, claims, onboarding files and contracts, extract their fields into structured data, check the values and pass the result to business systems. Documents that fail a check go to a person for review.

How accurate is AI document extraction?

Page reading is highly accurate on clean documents, with the best specialized models scoring about 96 on OmniDocBench. Field extraction is lower: GPT-4o got 73.3% of fields right on real invoices, and the best key-field score on the IDP Leaderboard is 91.1%. Old scans, handwriting, long tables and multi-page files are the weak spots.

Can we trust a model's confidence score to route documents to people?

Not on its own. In a study of GPT-4o on real invoices, token probabilities, stated confidence and agreement across runs all separated right from wrong answers poorly. Routing works when it combines failed checks, disagreement between two readings and page grounding, with thresholds set on a labeled sample of your own documents.

Do we still need OCR if we use a large language model?

Usually yes. Specialized parsing models beat frontier models on independent page-parsing benchmarks, and classic OCR pipelines are more reliable on handwriting and long documents. A common design uses a parser or OCR layer for text and positions, then a language model to extract fields, with the OCR positions used to confirm each value on the page.

How much does AI document processing cost?

Reading is cheap: managed OCR lists around $1.50 per 1,000 pages, and field extraction services tens of dollars per 1,000. The larger cost is people handling exceptions; Ardent Partners puts the average cost of processing an invoice at $9.84. Plan the budget around the share of documents that reach a person.

Will e-invoicing make invoice OCR unnecessary?

For domestic invoices in markets with mandates, largely yes. Germany, Belgium, Poland and France move most business invoices to structured formats by 2028, and the UK from 2029. Invoices from suppliers elsewhere, small suppliers, receipts and expense claims will keep arriving as images, as will trade documents, claims evidence and medical records.

Is AI document processing high-risk under the EU AI Act?

Extraction on its own usually is not. The Act treats a system that transforms unstructured data into structured data as a narrow procedural task. It becomes high-risk when it drives an Annex III decision such as creditworthiness, life or health insurance pricing or eligibility for public benefits, with those duties applying from December 2, 2027.

Sources

  1. OmniDocBench: benchmarking diverse PDF document parsingOpenDataLab, 2026
  2. ExtractConf: calibrated confidence for LLM document extractionarXiv, June 2026
  3. IDP Leaderboard: task detailsIDP Leaderboard (sponsored by Nanonets)
  4. Do VLMs Read or Rewrite? Faithfulness in OCRarXiv, 2026
  5. KIE-HVQA: OCR hallucination in degraded documentsarXiv, NeurIPS 2025
  6. Advancing Invoice Document Processing at Uber using GenAIUber Engineering, April 2025
  7. Rocket Close transforms mortgage document processing with Amazon Bedrock and Amazon TextractAWS Machine Learning Blog, April 2026
  8. When the storm clears, so should the claim queueAllianz, November 3, 2025
  9. FAQ: Einführung der E-RechnungGerman Federal Ministry of Finance
  10. Belgium's 2026 e-invoicing regulations explainedVertex
  11. Poland announces new timeline for mandatory e-invoicingEY
  12. France confirms September 1, 2026 e-invoicing deadlineSovos
  13. AI in accounts payable: the metrics that matterTungsten Automation, April 15, 2025
  14. olmOCR and olmOCR-BenchAllen Institute for AI, 2025
  15. Mistral OCRMistral AI, March 6, 2025
  16. ExtractBench: schema-guided enterprise document extractionarXiv, July 2026
  17. DISCO: OCR pipelines and vision-language models on document understandingarXiv, March 2026
  18. Xerox copier flaw means dodgy numbers and dangerous designsThe Register, August 6, 2013
  19. Structured OutputsOpenAI API documentation
  20. Structured outputsClaude Developer Platform documentation
  21. Three-way matching policiesMicrosoft Learn, Dynamics 365 Finance
  22. What is Azure Content Understanding?Microsoft Learn
  23. Bedrock Data AutomationAWS documentation
  24. HSBC goes proprietary on AI trade doc checkingGlobal Trade Review, 2026
  25. AI performs over three million shipping tasksC.H. Robinson, April 16, 2025
  26. Whitemail insights and vulnerability scanner: algorithmic transparency recordGOV.UK, 2025
  27. Home Office to expand AI use in asylum decision-makingElectronic Immigration Network, 2025
  28. Paper return processing: digitization resultsTIGTA, February 2026
  29. How Cigna saves millions by having its doctors reject claims without reading themProPublica, March 25, 2023
  30. Court partially grants, partially denies Cigna's motion to dismiss AI claims review caseDigital Healthcare Law, May 14, 2025
  31. Federal court orders broad discovery against UHC in AI coverage denial lawsuitArentFox Schiff, March 2026
  32. Regulation (EU) 2024/1689, the AI ActEUR-Lex
  33. Regulation (EU) 2026/1744, Digital Omnibus on AIEUR-Lex, July 2026
  34. CJEU rules that a credit score constitutes automated decision-making under the GDPRA&O Shearman, 2023
  35. SB-1120 Health care coverage: utilization reviewCalifornia Legislative Information, 2024
  36. Is AI WISeR? CMS models AI-based prior authorization in six statesDinsmore, 2025
  37. CMS-0057-F: Interoperability and Prior Authorization final ruleFirely
  38. Digital Personal Data Protection Rules, 2025: phased commencementPress Information Bureau, November 2025
  39. Regulation (EU) 2024/1624 on preventing the use of the financial system for money launderingEUR-Lex
  40. Revised time limit for e-invoice reporting for businesses with AATO of ₹10 crores and aboveGSTN
  41. Promoting electronic invoicing across UK businesses and the public sector: consultation responseGOV.UK, November 26, 2025
  42. VAT in the Digital Age (ViDA)European Commission
  43. eBL adoption doubles to 5% but barriers to digitisation remain, DCSA findsGlobal Trade Review, 2025
  44. Document AI pricingGoogle Cloud
  45. Amazon Textract pricingAWS
  46. olmOCR: unlocking trillions of tokens in PDFs with vision language modelsarXiv, 2025
  47. State of ePayables: AP benchmarks and best-in-class performanceArdent Partners, January 2026
  48. Marker: commercial usageDatalab on GitHub
  49. MinerU licenseOpenDataLab on GitHub
  50. dots.ocr license agreementrednote-hilab on Hugging Face
  51. Data, privacy and security for Document IntelligenceMicrosoft Learn
  52. Amazon Textract FAQsAWS
  53. Data controls in the OpenAI platformOpenAI API documentation
  54. API and data retentionClaude Developer Platform documentation

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 5, 2026. No client data appears in our insights.

Read next

All insights
  • On one plinth, a small glass prototype sits on a workbench at the front and drums of live data wait behind it; both run along lit conduits into a glowing go-live gate under a shield, and once it passes them a lit production tower rises at full height under a beam of light.

    Economics and buying Playbook

    AI Pilot to Production: How to Rescue, Restart, Buy or Stop a Stalled Pilot

    For CTOs, CIOs and heads of AI with a pilot that worked in the demo and has not been signed off to run for real, deciding whether to rescue it, rebuild it, buy instead or stop.

    22 min read

  • Your own cases, stored in a golden-set archive, run through your system to a release gate whose board shows every segment against a threshold its owner signed in advance; one segment falls short and the gate holds, then the rerun clears the line and the release crosses a bridge to production.

    LLM and RAG engineering Guide

    LLM and AI Agent Evaluation: How to Prove a System Is Ready to Ship

    For CTOs, heads of AI and risk owners deciding whether an LLM application, RAG system or AI agent is ready to leave the pilot, and what evidence should back that decision.

    16 min read

  • An AI agent's road runs into a policy hall that classes every action, and reads and reversible changes run on their own. The irreversible road alone leaves the plinth, across a bridge to the outside world, and a lit gate holds it until a named person beside it approves.

    AI agents Playbook

    AI Agent Guardrails: 8 Controls That Hold Up in Production

    For CTOs and CISOs deciding what an AI agent may do on its own before it touches customers, money or production.

    13 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.