LLM and RAG engineering Playbook
Enterprise RAG in Production: Why Pilots Stall and What Fixes Them
A stalled RAG pilot usually has a retrieval and data problem that the demo never exposed. This playbook sets out the seven ways RAG pilots fail in production, the fix for each, and the tests that show a system is ready for real users.
For CTOs deciding whether a RAG pilot that impressed in the demo can be trusted with real users, real permissions and real data.
The short answer
Enterprise RAG pilots stall because the demo tested a model on a small, open document set, and production tests retrieval across permissioned sources that change every day. The fixes sit in retrieval: hybrid keyword and vector search with a reranker, chunks that follow document structure, filters on each user's entitlements at query time, change-driven indexing with verified deletion, and evaluation that scores retrieval separately from generation.
Key takeaways
- Fix retrieval before you change the model: run keyword and vector search together, rerank the merged results, and cut chunks along the document's own structure.
- Filter by the user's entitlements inside the search, at query time, and never index a source whose access controls you cannot read.
- Design deletion from the start. An erasure request has to reach chunks, vectors, caches and logs, and researchers have recovered exact text from embeddings.9
- Call the system of record for live values such as balances and order status; retrieve text only for knowledge written to be read.
- Score retrieval and generation separately. Low recall means the answer never reached the model, and no prompt will fix it.
A RAG demo is built to impress: a few hundred clean documents, questions the team already knows the answers to, and a reviewer who reads the first paragraph. Production reverses every condition. The corpus spans every system the company runs, most of it under permissions; documents change and get deleted; and the people asking know when an answer is wrong. Retrieval-augmented generation was first described in 2020 as a way to give a language model an external memory it could cite and update,1 and it lives or dies on the retrieval half of its name.
- Nearly 1 in 3respondents in McKinsey's 2025 survey reported consequences from AI inaccuracy, the most reported AI risk2
- 17 to 33%of answers from commercial RAG legal research tools were hallucinated in a 2024 Stanford-led study3
- 60%of AI projects unsupported by AI-ready data that Gartner expects organizations to abandon through 20264
Why enterprise RAG pilots stall in production
Enterprise RAG pilots stall when the system meets the conditions the demo left out: exact terms, permissions, change, deletion and questions the documents cannot answer. In McKinsey's 2025 survey of 1,993 participants, inaccuracy was the AI risk respondents most often reported consequences from, and one of two risks that most respondents say their organizations are working to mitigate.2 In a July 2024 Gartner survey of 1,203 data management leaders, 63% of organizations lacked, or were unsure they had, the right data management practices for AI; Gartner lists vector stores, chunking and embedding among the practices to add.4
The clearest public evidence comes from law, where answers can be checked. Researchers at Stanford and Yale ran a preregistered set of more than 200 legal questions through three commercial RAG research tools from providers that had promoted RAG as avoiding or eliminating hallucinations. The tools hallucinated on 17 to 33% of answers, against 43% for a general-purpose chatbot with no retrieval.3 When the authors traced causes, poor retrieval contributed to 20 to 47% of the hallucinated answers depending on the tool, and citing an inapplicable source, such as an overruled case or the wrong jurisdiction, to 23 to 38%.3 Those are index, ranking and metadata problems. In the systems we build and review, the same pattern recurs: the model faithfully summarizes the wrong passage.
The seven production failure modes
A RAG pilot fails in production in seven recurring ways, and each has a symptom users report, a cause in the pipeline and a fix an engineer can schedule.
| Failure mode | What users see | Cause | Fix |
|---|---|---|---|
| Exact terms missed | A query for clause 14.3 or part AX-220 returns a loosely related passage | Vector search alone blurs identifiers, codes and names | Hybrid keyword and vector search, then a reranker |
| Broken context | The answer states a rule and drops its exception | Fixed-size chunks split sections, tables and conditions | Chunk by document structure, with the section path on every chunk |
| Permission leak | A user sees a salary band or another client's file | The index was built with a service account and stores no access controls | Filter by the user's entitlements inside the search, at query time |
| Stale or superseded answer | The assistant cites last year's policy | Batch reindexing; every version indexed as equal | Change-driven ingestion, version metadata, current version by default |
| Deleted data resurfaces | A record erased at the source still appears in answers | Deletes never reach chunks, vectors or caches | Lineage from document to chunk and a verified deletion pipeline |
| Wrong live figures | An order status or balance is days out of date | System-of-record data was exported and embedded as text | Call the system as a tool, with the user's authority |
| Invented answers | A fluent answer to a question the documents never address | No abstain path; uncited claims reach the user | Citations checked in code and a measured "I don't know" |
Six of the seven sit in retrieval and data. A larger model fixes none of them, which is why a stalled pilot needs a pipeline review before it needs a new model.
Fix retrieval first: hybrid search, reranking and chunking by structure
Retrieval quality sets the ceiling on answer quality, so fix it before touching prompts or models: combine keyword and vector search, rerank the merged candidates, and give the model a few passages cut along the document's own structure. The BEIR benchmark tested ten retrieval systems across 18 datasets outside their training domain and found the classic keyword method, BM25, a strong baseline, dense vector models often behind other approaches, and reranking models best on average, at a higher compute cost.5 Enterprise text, full of product codes, clause numbers, acronyms and names, is out of domain for any off-the-shelf embedding model; keyword search matches those terms exactly, and vectors blur them.
Run both searches, merge the two ranked lists with reciprocal rank fusion, and let a cross-encoder reranker order the top few dozen candidates. Then pass the model fewer passages. More context feels safer and performs worse: a 2023 study found that language models often use information best when it sits at the beginning or end of the input, and significantly worse when it sits in the middle of a long context, even for models built for long inputs.6 Keep the handful of passages that clear a relevance threshold, and put the strongest first.
Chunk by structure. Fixed-size windows cut a clause from its exception and a table from its header row. Split on the document's own headings, sections, list items and table rows. Attach the title, section path, effective date, owner and source link to every chunk, and keep a pointer to the parent section so the generator can see the surrounding rule when a small chunk matches. Scanned PDFs, slide decks and spreadsheets each need their own parser, and in the systems we build, parsing decides more of the final quality than the embedding model does.
Permission-aware retrieval: filter by the user's entitlements at query time
Permission-aware retrieval means every search runs as the person asking, and only passages that person could open in the source system are eligible to reach the model. OWASP's 2025 Top 10 for LLM applications lists vector and embedding weaknesses as a risk in its own right, naming unauthorized access through misaligned access controls and leaks between users who share a vector database, and it recommends permission-aware vector stores.7 The common mistake: a pilot indexes a document library with a service account that can read everything, and the assistant becomes a search box over every file that account could open. Filtering the answer afterwards is too late, because the restricted text is already in the model's context.
- Identify the user Resolve the signed-in user's groups and roles, including nested groups, from the OAuth or OpenID Connect token your identity provider issues.
- Filter inside the search Apply the entitlement filter within both the keyword and the vector query, so restricted chunks never become candidates. Filtering after retrieval returns thin results and wastes the top k.
- Re-check the most sensitive sources For HR, legal and board material, confirm access against the source system at query time, because the access list copied into the index can be hours old.
- Log the retrieval Record which passages were retrieved for whom, under which policy version, so an access question can be answered from the record.
Never index a source whose access controls you cannot read; if a system cannot export them, leave it out or index it only for a group allowed to see all of it. Sync permission changes on a stated interval and treat that interval as a security commitment, because a revoked user keeps access until the next sync. Key any answer cache by entitlement set, or one user's cached answer becomes another user's leak. In a multi-tenant system, keep each client in its own index or namespace.
Freshness, versioning and deletion
A production index must reflect each source within a stated time, know which version of a document is current, and forget what the source deletes. Nightly full reindexing is a demo habit. Drive ingestion from change events or modification timestamps in each source, hash every chunk so unchanged text is not embedded again, and publish a freshness target per source: minutes for the ticketing system, a day for the policy library. Give every chunk an effective date and a superseded-by link, and retrieve the current version by default. The inapplicable-authority errors in the Stanford study, such as citing an overruled case, are the legal form of this failure.3
Deletion has to be designed in from the start. Under Article 17 of the GDPR, a person can require a controller to erase their personal data without undue delay where one of the listed grounds applies, and Article 12 requires the controller to report the action taken within one month of the request, extendable by two further months for complex or numerous requests.8 In a RAG system that data lives in far more places than the source.
What a deletion request has to reach
- The source record, and the document-to-chunk lineage table that finds everything derived from it.
- Chunks in the keyword index and vectors in the vector index, including replicas.
- Cached answers, summaries, conversation logs and traces that stored its text.
- Evaluation sets and any training data built from real questions.
- Backups, on the schedule your retention policy states.
- A verification query that confirms nothing comes back, and a log entry that records the check.
Ask your vector store's supplier two questions: when a deleted vector stops being returned, and when it is physically removed. Write both answers into the retention policy. Fine-tuning is harder: a fact learned into model weights has no row to delete, and removing it reliably means retraining.
Retrieve the text or call the system
Retrieve text when the answer lives in prose written to be read; call the system of record when the answer is a live value, a calculation over records, or depends on row-level permissions. Order status, account balances, inventory, entitlements and case history change by the minute and already sit behind an API with its own authorization model. Exporting them to text and embedding the result creates a stale copy with the permissions stripped out. Give the model a tool instead: a typed call to the system's API, executed with the user's authority, returning structured data the answer can cite.
Retrieve
Search an index of documents
- Policies, contracts, manuals, research and ticket histories
- The answer is a passage a person could read and cite
- Tolerates a freshness lag of minutes to a day
- Permissions copied into the index and filtered at query time
Call the system
Query the system of record as a tool
- Orders, balances, stock levels, entitlements and case status
- The answer is a current value or a calculation over records
- Needs the value as it stands now
- Permissions enforced by the system's own API
The Model Context Protocol is an open protocol that gives models a standard way to reach those tools and data sources; its current specification is dated July 28, 2026.10 RAG and MCP work together: a retrieval service can itself be offered to the model as an MCP tool, next to tools that query ServiceNow, Salesforce or SAP. Our guide to MCP and A2A in the enterprise covers how to govern those connections. When the model decides for itself when to search and which tool to call, the pattern is called agentic RAG.
RAG vs fine-tuning vs long context
Use RAG for knowledge that changes, must be cited or depends on who is asking; use fine-tuning to change how a model behaves; use long context for one-off work over a small, known set of documents.
| Question | RAG | Fine-tuning | Long context |
|---|---|---|---|
| What it changes | What the model reads when it answers | The model's weights, and so its behavior | How much the model reads in one call |
| Best for | Facts that change, cited answers, per-user access | Output format, tone, domain vocabulary, classification | Analyzing a contract pack or a filing someone hands it |
| Freshness | As current as the index | Frozen at the last training run | Current for that call |
| Permissions and deletion | Filter per user; delete from the index | No per-user control; removing learned data means retraining | Nothing indexed; only what the caller passes in |
| Cost driver | Retrieval infrastructure plus a few passages per answer | Training runs, evaluation and hosting a custom model | Every token of every document, on every call |
A 2023 comparison of knowledge injection methods found that RAG consistently outperformed unsupervised fine-tuning, both for knowledge the model had seen in training and for entirely new facts, and that models struggle to learn new facts through unsupervised fine-tuning.11 Long context has improved without removing the position problem. On the NoLiMa benchmark, published in 2025, 11 of 13 models that claim at least 128K tokens of context fell below half their short-context score at 32K tokens, and even one of the strongest dropped from 99.3% to 69.7%.12 Long context is a good way to read one contract. It is a poor way to search ten thousand.
Evaluate retrieval separately from generation
Score retrieval and generation separately, because an end-to-end score cannot tell you which half to fix. Build the test set from real questions in the pilot's logs, each labeled with the passages that answer it, and add two kinds the demo never had: questions the corpus cannot answer, and questions a given test user is not permitted to see answered.
| Stage | Metric | What it measures | If it is low |
|---|---|---|---|
| Retrieval | Recall@k | Share of questions where a correct passage is in the top k | Fix search: hybrid queries, chunking, metadata |
| Retrieval | Precision@k | Share of retrieved passages that are relevant | Fix ranking: reranker, threshold, fewer passages |
| Retrieval | Permission leak rate | Restricted passages retrieved for users without access; the target is zero | Stop the release and fix the entitlement filter |
| Generation | Faithfulness | Share of claims in the answer supported by the retrieved passages | Fix the prompt, the citation check or the model |
| Generation | Answer relevance | Whether the answer addresses the question asked | Fix the prompt or the query handling |
| Generation | Abstention accuracy | Unanswerable questions correctly declined, and answerable ones answered | Tune the abstain threshold |
Low recall means the answer never reached the model, and no prompt will fix it. High recall with low faithfulness means the model had the right passage and drifted from it. Our guide to LLM and AI agent evaluation covers building the graded set and wiring it into your release pipeline.
Citations and "I don't know"
Every claim in a production answer should cite the passage it came from, and the system should say it does not know when no passage supports an answer. Enforce both in code. Ask the model to cite passage ids, then check that every cited id was in the retrieved set and that any quoted text appears in that passage; strip or flag a claim that fails. Link each citation to the source section, at the right version. Base the abstain rule on evidence: when no passage clears the reranker's threshold, the assistant says what it searched, says it found no answer, and routes the question to the owner of that content.
Tune abstention against both errors. In the Stanford study, a tool that tied for the lowest hallucination rate answered accurately on only 20% of queries, because 63% of its answers were refusals or lacked grounding.3 A system that rarely invents and rarely helps will stall as surely as one that invents. Track both rates, and send every unanswered question to the content owners as a gap report.
What drives the cost per answer
The cost of an answer is driven by the tokens sent to the model, the number of model and retrieval calls each answer takes, and the ingestion work spread across all answers. Model it before scale with a formula like this one:
cost per call = input tokens x input price + output tokens x output price
(input = instructions + passages + conversation history)
cost per answer = query embedding + keyword and vector search
+ candidates reranked x rerank price
+ calls per answer x cost per call (1 for single-shot RAG, more for agentic)
+ monthly ingestion and index cost / monthly answers
+ review minutes per answer x reviewer costFour levers move it. Passing five strong passages instead of twenty cuts input tokens and keeps the relevant one out of the middle of a long context. Caching answers to repeated questions cuts calls, provided the cache is keyed by entitlement set. Routing simple lookups to a smaller model lowers the price per call. Agentic retrieval multiplies calls per answer, so reserve it for the questions that need several steps.
What to ask before an enterprise RAG system goes live
Ask for a measured answer to each of these seven questions before go-live; any question without one marks work still to do.
Seven questions for the go-live review
- What are recall@k and faithfulness on our graded set, and which drop blocks a release?
- Does every search run as the signed-in user, and what is the measured permission leak rate?
- What is the freshness target for each source, and who is alerted when it is missed?
- When a source record is deleted, where is the proof that it is gone from every copy?
- Which questions go to a system of record instead of the index?
- What does the assistant say when it cannot find an answer, and who receives the gap?
- What does an answer cost today, and at ten times the volume?
This is the core of our AI development work: when we take over a stalled pilot, the first step is a retrieval and permission review against a graded set, so the fix goes where the failure is.
Questions leaders ask
Does RAG eliminate hallucinations?
No. Retrieval reduces hallucinations and leaves a residue. In a 2024 Stanford-led study, commercial RAG legal research tools hallucinated on 17 to 33% of answers, against 43% for a general-purpose chatbot without retrieval.3 Many of the remaining errors traced to retrieving the wrong or an inapplicable source, so the controls are better retrieval, citations checked in code, and an abstain path when no passage supports an answer.
Is RAG better than fine-tuning?
For adding knowledge, usually yes. A 2023 study found RAG consistently outperformed unsupervised fine-tuning for both familiar and new facts.11 RAG also keeps answers current, cites sources, respects per-user permissions and supports deletion, none of which model weights can do. Fine-tuning is the better tool for changing behavior: output format, tone, domain vocabulary or a classification task. The two combine well, with a fine-tuned model answering over retrieved passages.
What is the difference between RAG and MCP?
RAG is a pattern: search a corpus and give the model the passages that answer the question. MCP is an open protocol for connecting models to tools and data sources, and its current specification is dated July 28, 2026.10 The two combine. A retrieval service can be offered to a model as an MCP tool, next to tools that query live systems such as an order database or a ticketing system.
What is agentic RAG?
Agentic RAG lets the model decide when to search, which source or tool to use, and whether to search again after reading the first results. It handles multi-step questions that one search cannot, such as checking a contract against the policy it must follow. It costs more model calls per answer and adds latency, and every tool it can call needs the same permission checks as the retrieval itself.
What is the difference between enterprise search and RAG?
Enterprise search returns a ranked list of documents and leaves the reading to the person. RAG uses the same retrieval step, then has a model read the top passages and write an answer with citations. RAG therefore inherits every weakness of the search beneath it, and adds one more: a fluent answer hides a bad result that a list of links would have exposed.
Will long context windows replace RAG?
Not for enterprise knowledge. Models that accept very long inputs still lose accuracy as the input grows: in a 2025 benchmark, 11 of 13 long-context models fell below half their short-context score at 32K tokens.12 Long context also pays for every token on every call and does nothing to filter by user permissions. It suits reading one known set of documents; retrieval remains the way to search a company's knowledge.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksLewis et al., NeurIPS 2020 (arXiv 2005.11401, May 2020)
- The state of AI in 2025: Agents, innovation, and transformationMcKinsey & Company, November 5, 2025
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh, Surani, Dahl, Suzgun, Manning and Ho, Stanford University and Yale University, arXiv 2405.20362, May 2024
- Lack of AI-Ready Data Puts AI Projects at RiskGartner, February 26, 2025
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval ModelsThakur et al., NeurIPS 2021 Datasets and Benchmarks Track (arXiv 2104.08663)
- Lost in the Middle: How Language Models Use Long ContextsLiu et al., Transactions of the Association for Computational Linguistics (arXiv 2307.03172, July 2023)
- LLM08:2025 Vector and Embedding WeaknessesOWASP Top 10 for LLM Applications 2025
- Regulation (EU) 2016/679, the General Data Protection Regulation (Articles 12 and 17)Official Journal of the European Union, L 119, May 4, 2016
- Text Embeddings Reveal (Almost) As Much As TextMorris et al., EMNLP 2023 (arXiv 2310.06816)
- Model Context Protocol Specification, version 2026-07-28Model Context Protocol, July 28, 2026
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsOvadia et al., arXiv 2312.05934, December 2023
- NoLiMa: Long-Context Evaluation Beyond Literal MatchingModarressi et al., ICML 2025 (arXiv 2502.05167)
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on September 26, 2026. No client data appears in our insights.