LLM and RAG engineering Playbook

Enterprise RAG in Production: Why Pilots Stall and What Fixes Them

A stalled RAG pilot usually has a retrieval and data problem that the demo never exposed. This playbook sets out the seven ways RAG pilots fail in production, the fix for each, and the tests that show a system is ready for real users.

For CTOs deciding whether a RAG pilot that impressed in the demo can be trusted with real users, real permissions and real data.

Published
Reviewed
Reading time
16 min

The short answer

Enterprise RAG pilots stall because the demo tested a model on a small, open document set, and production tests retrieval across permissioned sources that change every day. The fixes sit in retrieval: hybrid keyword and vector search with a reranker, chunks that follow document structure, filters on each user's entitlements at query time, change-driven indexing with verified deletion, and evaluation that scores retrieval separately from generation.

Key takeaways

  • Fix retrieval before you change the model: run keyword and vector search together, rerank the merged results, and cut chunks along the document's own structure.
  • Filter by the user's entitlements inside the search, at query time, and never index a source whose access controls you cannot read.
  • Design deletion from the start. An erasure request has to reach chunks, vectors, caches and logs, and researchers have recovered exact text from embeddings.9
  • Call the system of record for live values such as balances and order status; retrieve text only for knowledge written to be read.
  • Score retrieval and generation separately. Low recall means the answer never reached the model, and no prompt will fix it.

A RAG demo is built to impress: a few hundred clean documents, questions the team already knows the answers to, and a reviewer who reads the first paragraph. Production reverses every condition. The corpus spans every system the company runs, most of it under permissions; documents change and get deleted; and the people asking know when an answer is wrong. Retrieval-augmented generation was first described in 2020 as a way to give a language model an external memory it could cite and update,1 and it lives or dies on the retrieval half of its name.

  • Nearly 1 in 3respondents in McKinsey's 2025 survey reported consequences from AI inaccuracy, the most reported AI risk2
  • 17 to 33%of answers from commercial RAG legal research tools were hallucinated in a 2024 Stanford-led study3
  • 60%of AI projects unsupported by AI-ready data that Gartner expects organizations to abandon through 20264

Why enterprise RAG pilots stall in production

Enterprise RAG pilots stall when the system meets the conditions the demo left out: exact terms, permissions, change, deletion and questions the documents cannot answer. In McKinsey's 2025 survey of 1,993 participants, inaccuracy was the AI risk respondents most often reported consequences from, and one of two risks that most respondents say their organizations are working to mitigate.2 In a July 2024 Gartner survey of 1,203 data management leaders, 63% of organizations lacked, or were unsure they had, the right data management practices for AI; Gartner lists vector stores, chunking and embedding among the practices to add.4

The clearest public evidence comes from law, where answers can be checked. Researchers at Stanford and Yale ran a preregistered set of more than 200 legal questions through three commercial RAG research tools from providers that had promoted RAG as avoiding or eliminating hallucinations. The tools hallucinated on 17 to 33% of answers, against 43% for a general-purpose chatbot with no retrieval.3 When the authors traced causes, poor retrieval contributed to 20 to 47% of the hallucinated answers depending on the tool, and citing an inapplicable source, such as an overruled case or the wrong jurisdiction, to 23 to 38%.3 Those are index, ranking and metadata problems. In the systems we build and review, the same pattern recurs: the model faithfully summarizes the wrong passage.

Exhibit 1Retrieval reduced hallucinations and did not remove them
  • General-purpose chatbot, no retrieval43%
  • RAG legal research tool A17%
  • RAG legal research tool B33%
  • RAG legal research tool C17%
Share of answers with a hallucination over more than 200 preregistered legal queries, evaluated by April 2024; at least one vendor has since shipped a newer version. Tool C's low rate came with refusals or ungrounded answers on 63% of queries. Source: [3]

The seven production failure modes

A RAG pilot fails in production in seven recurring ways, and each has a symptom users report, a cause in the pipeline and a fix an engineer can schedule.

Failure modeWhat users seeCauseFix
Exact terms missedA query for clause 14.3 or part AX-220 returns a loosely related passageVector search alone blurs identifiers, codes and namesHybrid keyword and vector search, then a reranker
Broken contextThe answer states a rule and drops its exceptionFixed-size chunks split sections, tables and conditionsChunk by document structure, with the section path on every chunk
Permission leakA user sees a salary band or another client's fileThe index was built with a service account and stores no access controlsFilter by the user's entitlements inside the search, at query time
Stale or superseded answerThe assistant cites last year's policyBatch reindexing; every version indexed as equalChange-driven ingestion, version metadata, current version by default
Deleted data resurfacesA record erased at the source still appears in answersDeletes never reach chunks, vectors or cachesLineage from document to chunk and a verified deletion pipeline
Wrong live figuresAn order status or balance is days out of dateSystem-of-record data was exported and embedded as textCall the system as a tool, with the user's authority
Invented answersA fluent answer to a question the documents never addressNo abstain path; uncited claims reach the userCitations checked in code and a measured "I don't know"
Symptoms reach the help desk. Causes and fixes are where the engineering time goes.

Six of the seven sit in retrieval and data. A larger model fixes none of them, which is why a stalled pilot needs a pipeline review before it needs a new model.

Fix retrieval first: hybrid search, reranking and chunking by structure

Retrieval quality sets the ceiling on answer quality, so fix it before touching prompts or models: combine keyword and vector search, rerank the merged candidates, and give the model a few passages cut along the document's own structure. The BEIR benchmark tested ten retrieval systems across 18 datasets outside their training domain and found the classic keyword method, BM25, a strong baseline, dense vector models often behind other approaches, and reranking models best on average, at a higher compute cost.5 Enterprise text, full of product codes, clause numbers, acronyms and names, is out of domain for any off-the-shelf embedding model; keyword search matches those terms exactly, and vectors blur them.

Run both searches, merge the two ranked lists with reciprocal rank fusion, and let a cross-encoder reranker order the top few dozen candidates. Then pass the model fewer passages. More context feels safer and performs worse: a 2023 study found that language models often use information best when it sits at the beginning or end of the input, and significantly worse when it sits in the middle of a long context, even for models built for long inputs.6 Keep the handful of passages that clear a relevance threshold, and put the strongest first.

Chunk by structure. Fixed-size windows cut a clause from its exception and a table from its header row. Split on the document's own headings, sections, list items and table rows. Attach the title, section path, effective date, owner and source link to every chunk, and keep a pointer to the parent section so the generator can see the surrounding rule when a small chunk matches. Scanned PDFs, slide decks and spreadsheets each need their own parser, and in the systems we build, parsing decides more of the final quality than the embedding model does.

Permission-aware retrieval: filter by the user's entitlements at query time

Permission-aware retrieval means every search runs as the person asking, and only passages that person could open in the source system are eligible to reach the model. OWASP's 2025 Top 10 for LLM applications lists vector and embedding weaknesses as a risk in its own right, naming unauthorized access through misaligned access controls and leaks between users who share a vector database, and it recommends permission-aware vector stores.7 The common mistake: a pilot indexes a document library with a service account that can read everything, and the assistant becomes a search box over every file that account could open. Filtering the answer afterwards is too late, because the restricted text is already in the model's context.

  1. Identify the user Resolve the signed-in user's groups and roles, including nested groups, from the OAuth or OpenID Connect token your identity provider issues.
  2. Filter inside the search Apply the entitlement filter within both the keyword and the vector query, so restricted chunks never become candidates. Filtering after retrieval returns thin results and wastes the top k.
  3. Re-check the most sensitive sources For HR, legal and board material, confirm access against the source system at query time, because the access list copied into the index can be hours old.
  4. Log the retrieval Record which passages were retrieved for whom, under which policy version, so an access question can be answered from the record.

Never index a source whose access controls you cannot read; if a system cannot export them, leave it out or index it only for a group allowed to see all of it. Sync permission changes on a stated interval and treat that interval as a security commitment, because a revoked user keeps access until the next sync. Key any answer cache by entitlement set, or one user's cached answer becomes another user's leak. In a multi-tenant system, keep each client in its own index or namespace.

Freshness, versioning and deletion

A production index must reflect each source within a stated time, know which version of a document is current, and forget what the source deletes. Nightly full reindexing is a demo habit. Drive ingestion from change events or modification timestamps in each source, hash every chunk so unchanged text is not embedded again, and publish a freshness target per source: minutes for the ticketing system, a day for the policy library. Give every chunk an effective date and a superseded-by link, and retrieve the current version by default. The inapplicable-authority errors in the Stanford study, such as citing an overruled case, are the legal form of this failure.3

Deletion has to be designed in from the start. Under Article 17 of the GDPR, a person can require a controller to erase their personal data without undue delay where one of the listed grounds applies, and Article 12 requires the controller to report the action taken within one month of the request, extendable by two further months for complex or numerous requests.8 In a RAG system that data lives in far more places than the source.

What a deletion request has to reach

  • The source record, and the document-to-chunk lineage table that finds everything derived from it.
  • Chunks in the keyword index and vectors in the vector index, including replicas.
  • Cached answers, summaries, conversation logs and traces that stored its text.
  • Evaluation sets and any training data built from real questions.
  • Backups, on the schedule your retention policy states.
  • A verification query that confirms nothing comes back, and a log entry that records the check.

Ask your vector store's supplier two questions: when a deleted vector stops being returned, and when it is physically removed. Write both answers into the retention policy. Fine-tuning is harder: a fact learned into model weights has no row to delete, and removing it reliably means retraining.

Retrieve the text or call the system

Retrieve text when the answer lives in prose written to be read; call the system of record when the answer is a live value, a calculation over records, or depends on row-level permissions. Order status, account balances, inventory, entitlements and case history change by the minute and already sit behind an API with its own authorization model. Exporting them to text and embedding the result creates a stale copy with the permissions stripped out. Give the model a tool instead: a typed call to the system's API, executed with the user's authority, returning structured data the answer can cite.

Exhibit 2Two ways to answer from enterprise data

Retrieve

Search an index of documents

  • Policies, contracts, manuals, research and ticket histories
  • The answer is a passage a person could read and cite
  • Tolerates a freshness lag of minutes to a day
  • Permissions copied into the index and filtered at query time

Call the system

Query the system of record as a tool

  • Orders, balances, stock levels, entitlements and case status
  • The answer is a current value or a calculation over records
  • Needs the value as it stands now
  • Permissions enforced by the system's own API
A production assistant usually needs both, and the router that chooses between them is part of the design.

The Model Context Protocol is an open protocol that gives models a standard way to reach those tools and data sources; its current specification is dated July 28, 2026.10 RAG and MCP work together: a retrieval service can itself be offered to the model as an MCP tool, next to tools that query ServiceNow, Salesforce or SAP. Our guide to MCP and A2A in the enterprise covers how to govern those connections. When the model decides for itself when to search and which tool to call, the pattern is called agentic RAG.

RAG vs fine-tuning vs long context

Use RAG for knowledge that changes, must be cited or depends on who is asking; use fine-tuning to change how a model behaves; use long context for one-off work over a small, known set of documents.

QuestionRAGFine-tuningLong context
What it changesWhat the model reads when it answersThe model's weights, and so its behaviorHow much the model reads in one call
Best forFacts that change, cited answers, per-user accessOutput format, tone, domain vocabulary, classificationAnalyzing a contract pack or a filing someone hands it
FreshnessAs current as the indexFrozen at the last training runCurrent for that call
Permissions and deletionFilter per user; delete from the indexNo per-user control; removing learned data means retrainingNothing indexed; only what the caller passes in
Cost driverRetrieval infrastructure plus a few passages per answerTraining runs, evaluation and hosting a custom modelEvery token of every document, on every call
The three combine: a fine-tuned model can answer over retrieved passages.

A 2023 comparison of knowledge injection methods found that RAG consistently outperformed unsupervised fine-tuning, both for knowledge the model had seen in training and for entirely new facts, and that models struggle to learn new facts through unsupervised fine-tuning.11 Long context has improved without removing the position problem. On the NoLiMa benchmark, published in 2025, 11 of 13 models that claim at least 128K tokens of context fell below half their short-context score at 32K tokens, and even one of the strongest dropped from 99.3% to 69.7%.12 Long context is a good way to read one contract. It is a poor way to search ten thousand.

Evaluate retrieval separately from generation

Score retrieval and generation separately, because an end-to-end score cannot tell you which half to fix. Build the test set from real questions in the pilot's logs, each labeled with the passages that answer it, and add two kinds the demo never had: questions the corpus cannot answer, and questions a given test user is not permitted to see answered.

StageMetricWhat it measuresIf it is low
RetrievalRecall@kShare of questions where a correct passage is in the top kFix search: hybrid queries, chunking, metadata
RetrievalPrecision@kShare of retrieved passages that are relevantFix ranking: reranker, threshold, fewer passages
RetrievalPermission leak rateRestricted passages retrieved for users without access; the target is zeroStop the release and fix the entitlement filter
GenerationFaithfulnessShare of claims in the answer supported by the retrieved passagesFix the prompt, the citation check or the model
GenerationAnswer relevanceWhether the answer addresses the question askedFix the prompt or the query handling
GenerationAbstention accuracyUnanswerable questions correctly declined, and answerable ones answeredTune the abstain threshold
Run the suite on every change to parsing, chunking, embeddings, ranking, prompts or models, and block the release when a metric falls.

Low recall means the answer never reached the model, and no prompt will fix it. High recall with low faithfulness means the model had the right passage and drifted from it. Our guide to LLM and AI agent evaluation covers building the graded set and wiring it into your release pipeline.

Citations and "I don't know"

Every claim in a production answer should cite the passage it came from, and the system should say it does not know when no passage supports an answer. Enforce both in code. Ask the model to cite passage ids, then check that every cited id was in the retrieved set and that any quoted text appears in that passage; strip or flag a claim that fails. Link each citation to the source section, at the right version. Base the abstain rule on evidence: when no passage clears the reranker's threshold, the assistant says what it searched, says it found no answer, and routes the question to the owner of that content.

Tune abstention against both errors. In the Stanford study, a tool that tied for the lowest hallucination rate answered accurately on only 20% of queries, because 63% of its answers were refusals or lacked grounding.3 A system that rarely invents and rarely helps will stall as surely as one that invents. Track both rates, and send every unanswered question to the content owners as a gap report.

What drives the cost per answer

The cost of an answer is driven by the tokens sent to the model, the number of model and retrieval calls each answer takes, and the ingestion work spread across all answers. Model it before scale with a formula like this one:

cost per call   = input tokens x input price + output tokens x output price
                  (input = instructions + passages + conversation history)

cost per answer = query embedding + keyword and vector search
                + candidates reranked x rerank price
                + calls per answer x cost per call        (1 for single-shot RAG, more for agentic)
                + monthly ingestion and index cost / monthly answers
                + review minutes per answer x reviewer cost
Every term is a design choice. None is fixed by the technology.

Four levers move it. Passing five strong passages instead of twenty cuts input tokens and keeps the relevant one out of the middle of a long context. Caching answers to repeated questions cuts calls, provided the cache is keyed by entitlement set. Routing simple lookups to a smaller model lowers the price per call. Agentic retrieval multiplies calls per answer, so reserve it for the questions that need several steps.

What to ask before an enterprise RAG system goes live

Ask for a measured answer to each of these seven questions before go-live; any question without one marks work still to do.

Seven questions for the go-live review

  • What are recall@k and faithfulness on our graded set, and which drop blocks a release?
  • Does every search run as the signed-in user, and what is the measured permission leak rate?
  • What is the freshness target for each source, and who is alerted when it is missed?
  • When a source record is deleted, where is the proof that it is gone from every copy?
  • Which questions go to a system of record instead of the index?
  • What does the assistant say when it cannot find an answer, and who receives the gap?
  • What does an answer cost today, and at ten times the volume?

This is the core of our AI development work: when we take over a stalled pilot, the first step is a retrieval and permission review against a graded set, so the fix goes where the failure is.

Questions leaders ask

Does RAG eliminate hallucinations?

No. Retrieval reduces hallucinations and leaves a residue. In a 2024 Stanford-led study, commercial RAG legal research tools hallucinated on 17 to 33% of answers, against 43% for a general-purpose chatbot without retrieval.3 Many of the remaining errors traced to retrieving the wrong or an inapplicable source, so the controls are better retrieval, citations checked in code, and an abstain path when no passage supports an answer.

Is RAG better than fine-tuning?

For adding knowledge, usually yes. A 2023 study found RAG consistently outperformed unsupervised fine-tuning for both familiar and new facts.11 RAG also keeps answers current, cites sources, respects per-user permissions and supports deletion, none of which model weights can do. Fine-tuning is the better tool for changing behavior: output format, tone, domain vocabulary or a classification task. The two combine well, with a fine-tuned model answering over retrieved passages.

What is the difference between RAG and MCP?

RAG is a pattern: search a corpus and give the model the passages that answer the question. MCP is an open protocol for connecting models to tools and data sources, and its current specification is dated July 28, 2026.10 The two combine. A retrieval service can be offered to a model as an MCP tool, next to tools that query live systems such as an order database or a ticketing system.

What is agentic RAG?

Agentic RAG lets the model decide when to search, which source or tool to use, and whether to search again after reading the first results. It handles multi-step questions that one search cannot, such as checking a contract against the policy it must follow. It costs more model calls per answer and adds latency, and every tool it can call needs the same permission checks as the retrieval itself.

What is the difference between enterprise search and RAG?

Enterprise search returns a ranked list of documents and leaves the reading to the person. RAG uses the same retrieval step, then has a model read the top passages and write an answer with citations. RAG therefore inherits every weakness of the search beneath it, and adds one more: a fluent answer hides a bad result that a list of links would have exposed.

Will long context windows replace RAG?

Not for enterprise knowledge. Models that accept very long inputs still lose accuracy as the input grows: in a 2025 benchmark, 11 of 13 long-context models fell below half their short-context score at 32K tokens.12 Long context also pays for every token on every call and does nothing to filter by user permissions. It suits reading one known set of documents; retrieval remains the way to search a company's knowledge.

Sources

  1. Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksLewis et al., NeurIPS 2020 (arXiv 2005.11401, May 2020)
  2. The state of AI in 2025: Agents, innovation, and transformationMcKinsey & Company, November 5, 2025
  3. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh, Surani, Dahl, Suzgun, Manning and Ho, Stanford University and Yale University, arXiv 2405.20362, May 2024
  4. Lack of AI-Ready Data Puts AI Projects at RiskGartner, February 26, 2025
  5. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval ModelsThakur et al., NeurIPS 2021 Datasets and Benchmarks Track (arXiv 2104.08663)
  6. Lost in the Middle: How Language Models Use Long ContextsLiu et al., Transactions of the Association for Computational Linguistics (arXiv 2307.03172, July 2023)
  7. LLM08:2025 Vector and Embedding WeaknessesOWASP Top 10 for LLM Applications 2025
  8. Regulation (EU) 2016/679, the General Data Protection Regulation (Articles 12 and 17)Official Journal of the European Union, L 119, May 4, 2016
  9. Text Embeddings Reveal (Almost) As Much As TextMorris et al., EMNLP 2023 (arXiv 2310.06816)
  10. Model Context Protocol Specification, version 2026-07-28Model Context Protocol, July 28, 2026
  11. Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsOvadia et al., arXiv 2312.05934, December 2023
  12. NoLiMa: Long-Context Evaluation Beyond Literal MatchingModarressi et al., ICML 2025 (arXiv 2502.05167)

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on September 26, 2026. No client data appears in our insights.

Read next

All insights
  • Your own cases, stored in a golden-set archive, run through your system to a release gate whose board shows every segment against a threshold its owner signed in advance; one segment falls short and the gate holds, then the rerun clears the line and the release crosses a bridge to production.

    LLM and RAG engineering Guide

    LLM and AI Agent Evaluation: How to Prove a System Is Ready to Ship

    For CTOs, heads of AI and risk owners deciding whether an LLM application, RAG system or AI agent is ready to leave the pilot, and what evidence should back that decision.

    16 min read

  • One walled estate of customer data has a gate for each AI use case, and each gate runs the same eight readiness checks: every lamp turns green for a churn model, which is built, while the freshness check turns red for a service agent, whose bar stays down and whose tower is still only drawn.

    Data foundations Checklist

    Is Your Data Ready for AI? A Use-Case Data Readiness Assessment

    For CTOs, chief data officers and CFOs deciding whether the data behind a proposed AI use case can carry it before the budget is committed.

    16 min read

  • On one plinth, your AI agent reaches an ERP, a CRM and a service desk through a single lit MCP gateway; a bridge in the air carries a task over A2A to a partner company's agent on its own island.

    AI agents Guide

    MCP and A2A in the Enterprise: How AI Agents Reach Your Systems

    For CTOs and CISOs deciding how AI agents will connect to SAP, Salesforce and the rest of the estate, and which systems to open to them first.

    16 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.