LLM and RAG engineering Guide
LLM and AI Agent Evaluation: How to Prove a System Is Ready to Ship
Public benchmarks tell you what a model can do on someone else's questions. Evaluation tells you whether your system does your work well enough to ship, and it rests on three things: a golden set built from your own cases, graders calibrated against your experts, and thresholds a named owner signs before the results are in.
For CTOs, heads of AI and risk owners deciding whether an LLM application, RAG system or AI agent is ready to leave the pilot, and what evidence should back that decision.
The short answer
LLM evaluation is the repeatable test that shows whether an LLM application, RAG system or AI agent does your work well enough to ship. It runs a golden set of your own historical cases, labeled by domain experts, on every change; scores outputs with code checks, calibrated LLM judges and expert review; and compares the results with release thresholds signed by a named business owner.
Key takeaways
- Leaderboards help you shortlist models and say little about your system. A release decision needs a golden set of your own cases, including the ones your experts argued about.5
- Size the golden set from the margin of error you can accept: at a 90% pass rate, 100 cases leave about 6 points of uncertainty either way, and 400 cases about 3.
- Treat an LLM judge as an instrument to calibrate. Published results range from 85% agreement with experts on pairwise votes to judges well behind human agreement on scores, so measure yours on your own cases.29
- Grade agents on end state, trajectory and consistency. On one 2024 benchmark, state-of-the-art agents succeeded on all eight attempts at a retail task less than 25% of the time.8
- Every gated metric needs a threshold set in advance and a named owner who signs it. The EU AI Act requires high-risk systems to be tested against prior defined metrics and probabilistic thresholds.10
A pilot passes its demo because someone chose the questions. Production chooses its own: the scanned invoice with a rotated page, the policy that changed last month, the customer who asks two things at once. Evaluation measures how often the system gets those right, and it is the gate between pilot and production. McKinsey's 2025 survey found that the organizations getting the most value from AI are more likely to have defined processes for deciding how and when model outputs need human validation to ensure accuracy.1 When a pilot stalls, the evaluation is the first thing we build.
- ~1 in 3respondents to McKinsey's 2025 global survey reported negative consequences from AI inaccuracy, the risk most often reported1
- 85%agreement between the strongest LLM judge and human experts on non-tied votes, against 81% among the experts themselves (2023)2
- 84%→51%one hosted model's accuracy on the same task, between its March and June 2023 versions3
Benchmarks versus evaluation
A benchmark scores a model on a public question set; an evaluation scores your whole system on your own cases, and only the evaluation can support a release decision. Leaderboards are useful for shortlisting models. They say little about your prompts, documents, tools or the cases your experts find hard.
Stanford's 2026 AI Index shows the gap. Top scores on a widely used coding benchmark rose from 60% to near 100% in one year, so the leaderboard stopped separating the leaders. Capability also stays uneven: a model has won gold at the International Mathematical Olympiad, yet the best model reads an analog clock correctly only 50.1% of the time.4 NIST's Generative AI Profile warns against extrapolating from narrow, anecdotal assessments and asks for performance demonstrated in conditions similar to deployment.5
Public benchmark
Someone else's questions
- A public question set any developer can tune against
- Measures the model alone, without your prompts, data or tools
- Saturates fast: one coding benchmark went from 60% to near 100% in a year
- Good for shortlisting models
Your evaluation
Your cases, your thresholds
- Real historical cases, labeled by your experts
- Measures the whole system, policy code included
- Grows with every production failure
- Decides whether a release ships
How to build a golden dataset from your own cases
Build a golden dataset from real cases in your own history, each holding the input, the context the system should use and the outcome a domain expert has confirmed as correct. It is the most durable asset in an LLM program: prompts, models and vendors change, and the golden set is how you judge each change.
- Sample from history Draw a year of real tickets, documents or transactions, stratified by case type, channel, language and value, so the set mirrors traffic and weights the cases that carry money or regulatory exposure.
- Add the hard cases on purpose Escalations, complaints, cases your experts argued about, requests the system must refuse, and adversarial inputs such as instructions hidden in a document. Random sampling under-represents the cases that cause incidents.
- Label with experts and a rubric Domain experts write the expected outcome and the reason against a written rubric; builders never label their own test. Two experts label a sample independently, and their agreement sets the ceiling for any automated grader.
- Hold out a slice Keep a portion nobody tunes prompts against, and run it only at release. A set the team has optimized against for months overstates quality.
- Version and refresh Record the set's version, label provenance and the policy version it reflects. Add every production failure, retire cases when policy changes and review the whole set each quarter.
Label quality decides what the set can prove. A NeurIPS study of ten widely used test sets estimated at least 3.3% label errors on average, and showed that a modest rise in label errors could reverse which of two models ranked higher.6 A set labeled in a hurry can carry more, and every wrong label hides a failure or invents one.
How many cases a golden set needs
Size the set from the margin of error you can accept. For a pass rate p on n independent cases, the 95% margin is about 1.96 × √(p × (1 − p) / n). At a 90% pass rate, that gives:
| Cases | Margin at a 90% pass rate | Typical use |
|---|---|---|
| 50 | ±8.3 points | Smoke test on every commit; catches large breaks |
| 100 | ±5.9 points | Early pilot; detects drops of about 10 points |
| 200 | ±4.2 points | One segment with its own threshold |
| 400 | ±2.9 points | Release gate that must catch drops of about 5 points |
| 1,000 | ±1.9 points | Comparing two strong models, or many segments |
Three adjustments follow. Each segment with its own threshold (a language, a product line, a document type) needs its own count. Related cases, such as several questions about one contract, are not independent; one statistical analysis of LLM evaluations found clustered standard errors more than three times the naive figure.7 And strict gates follow the rule of three: with zero failures in n cases, the true failure rate is below about 3/n at 95% confidence, so claiming under 1% errors on payment fields takes at least 300 clean cases.
What to measure for LLM applications, RAG systems and AI agents
Measure what each system is for: answers for an LLM application, retrieval and faithfulness to sources for a RAG system, actions and end states for an agent, and cost, latency and safety for all three.
| System type | What to measure | How it is measured |
|---|---|---|
| LLM application (extraction, classification, drafting) | Field accuracy; valid format; policy compliance; correct refusals | Code checks against labeled answers; schema validation; a calibrated judge for free text |
| RAG system | Retrieval recall; faithfulness of each claim to its source; answer correctness; citation accuracy; abstention when the sources hold no answer | Experts tag the passages that answer each question; a judge checks each claim against the retrieved text; unanswerable questions included on purpose |
| AI agent | Task success by end state; trajectory (tools, arguments, order); forbidden actions; escalation; consistency across runs | Compare the records left with the expected state; assert on the trace; repeat each task and report pass^k |
| All three | p95 latency; cost per completed task; harmful or leaked output; resistance to injected instructions | Timers and cost meters on every run; a red-team slice before each release |
Score a RAG system's retrieval and generation separately, because they fail separately. A wrong answer with the right passage retrieved is a generation problem; a wrong answer built on the wrong passages is a retrieval problem that no prompt change will fix. NIST's Generative AI Profile asks teams to verify sources and citations in outputs before deployment and in ongoing monitoring.5 More in Enterprise RAG in production.
How to evaluate an AI agent: outcome, trajectory and consistency
Grade an agent on what it did, how it did it and whether it does it every time. For the outcome, compare the state it left with the expected state; one widely cited agent benchmark compares the database at the end of each conversation with an annotated goal state.8 For the trajectory, assert on the trace: tools called, arguments, order, approvals requested and any forbidden action. An agent that issues the right refund after opening a record it had no reason to read passes the first test and fails the second.
Consistency is the measure pilots skip. The same benchmark introduced pass^k, the chance an agent succeeds on all k attempts at a task, and found in 2024 that state-of-the-art function-calling agents succeeded on fewer than half its tasks and scored below 25% on pass^8 in its retail domain.8 The arithmetic is unforgiving: 90% success per independent run means eight clean runs in a row only 43% of the time. Gate agents on pass^k; the controls that make the remaining failures safe are in AI agent guardrails that hold up in production.
LLM as a judge: how to calibrate it against human graders
Calibrate an LLM judge by running it on cases your experts have already labeled, and comparing its agreement with them to the experts' agreement with each other. The method rests on a 2023 study: the strongest judge tested agreed with human experts on 85% of non-tied votes, above the 81% among the humans themselves.2
The same study documented the biases. With the two answers swapped, the strongest judge gave a consistent verdict in only 65% of cases (position bias). An answer padded with a restated list beat the original with two of the three judges 91.3% of the time (verbosity bias). Judges also showed signs of favoring their own model's answers, which the authors could not confirm statistically. On math, the strongest judge accepted far fewer wrong answers when given the correct one.2
Later work is less optimistic. A 2025 study of thirteen judge models found only the best and largest reasonably aligned with humans, still well behind agreement between humans, with scores up to 5 points off and a lean toward leniency; high percent agreement could hide very different scores.9 The 2023 figure measures agreement on which of two answers is better, ties excluded, while the 2025 study compares absolute scores. In the systems we build, judges return pass or fail against a rubric, with a reason, because a binary verdict is easier to calibrate than a score.
How to calibrate an LLM judge
- Have two experts label the same 100 to 200 cases; their agreement is the judge's ceiling.
- Report the judge's Cohen's kappa against expert labels as well as raw agreement, per segment.
- Give the judge the rubric and the golden set's reference answer.
- Run pairwise comparisons in both orders and count a split verdict as a tie.
- Use a judge from a different model family than the system it grades.
- Pin the judge's version, and recalibrate when the judge, rubric or case mix changes.
- Keep experts reviewing a sample of production, so judge drift shows within weeks.
Release gates: who signs the threshold
Every metric in a release gate needs a threshold, the slice of the golden set it is measured on and a named business owner who signs it before testing starts. Engineers can measure a pass rate; they cannot decide what error rate a claims process or a payments run can bear. The EU AI Act requires high-risk systems to be tested against prior defined metrics and probabilistic thresholds,10 and NIST's Generative AI Profile asks that pre-deployment results go to those with release approval authority.5
| Metric | Example threshold | Measured on | Signed by |
|---|---|---|---|
| Critical-field accuracy (amounts, dates, bank details) | 99% or higher; zero errors on bank details | Invoice slice | Finance controller |
| Claims supported by the cited source | 98% or higher | Policy-question slice | Head of compliance |
| Out-of-scope requests handed to a person | 95% or higher | Out-of-scope and adversarial slice | Operations lead |
| Agent success across repeated runs | pass^4 of 90% or higher | Agent slice, four runs per case | Process owner |
| Harmful, non-compliant or leaked output | Zero | Red-team slice | CISO |
| Regression against the current release | No gated segment down beyond its margin | Full set, case by case | Engineering lead |
| Cost per completed task; p95 latency | Within the approved budget and target | Full set and shadow traffic | Product owner |
Two rules keep gates honest. Thresholds are fixed before results are seen, so nobody moves the line to fit the number. A failed gate means fix and rerun, or a signed exception from the owner with a reason and an expiry date in the release file.
Running LLM evaluation in CI on every change
Any change that can alter an output runs the golden set before it merges: a prompt, a model version, a retrieval setting, a tool definition, a policy file or a library upgrade. Run it in three tiers: a smoke slice of about 50 cases on every commit, the full set nightly and on every release candidate, and the red-team and hold-out slices before release. The build fails when a gated metric misses its threshold.
Compare the candidate with the current release case by case. A statistical analysis of LLM evaluations recommends inference on question-level paired differences when comparing two models, which strips out the variance from questions that are hard for both.7 Report the difference with its confidence interval, list every case that flipped from pass to fail, and store each run with its prompt, model, data and golden-set versions so it can be reproduced.
Online evaluation and drift in production
Online evaluation scores a sample of live traffic continuously, because inputs, source documents and even a hosted model change after release. A 2023 study of one widely used hosted model found its accuracy at identifying prime numbers fell from 84% to 51% between its March and June versions, and concluded that the same service can change substantially in a short time.3 Pin model versions where your provider allows it, and watch the numbers where it does not.
NIST's AI Risk Management Framework expects AI systems to be tested before deployment and regularly while in operation,11 and the EU AI Act requires providers of high-risk systems to monitor performance throughout the system's lifetime.10 Watch six signals:
- Calibrated judge scores on a random sample of live outputs.
- Human signals: edits, overrides, escalations and complaints.
- Input drift: topics, languages and document types the golden set lacks.
- Retrieval health: queries that return nothing or only weak passages.
- Agent behavior: steps per task, tool errors, retries and approval requests.
- Cost per completed task and p95 latency.
- Offline gateThe golden set and red-team slice run in CI against the release candidate.Every gated metric meets its signed threshold.
- ShadowThe candidate runs on live traffic beside the current system; its outputs are scored and none reaches a customer.Online scores match offline results within the margin.
- CanaryA small share of traffic goes to the candidate, with online evaluation and one-step rollback.User signals and sampled expert review hold for an agreed period.
- General releaseAll traffic moves over; online evaluation continues, and every failure becomes a golden case.Automatic rollback when a monitored metric crosses its alert line.
How to evaluate a model swap
Treat a model swap as a release: run the full golden set on the current and candidate models, compare them case by case, and switch only if every gated segment holds at acceptable cost and latency. Newer models are usually better on average and worse somewhere specific; Stanford's 2026 AI Index cites research finding that improving one responsible AI dimension, such as safety, can degrade another, such as accuracy.4 Give the candidate its own prompt tuning on the development slice before comparing the two on the hold-out.
| Dimension | What to compare | Switch only if |
|---|---|---|
| Quality | Paired, per-segment difference on the full golden set | No gated segment is worse beyond its margin |
| Regressions | Cases the current model passes and the candidate fails | An expert has read each one, and none breaks a gate |
| Safety | Red-team and adversarial slices | No new failures |
| Cost | Cost per completed task, with retries, longer outputs and judge calls | Within budget at forecast volume |
| Latency | p50 and p95 end to end, tool calls included | Within the response-time target |
Evaluation evidence for auditors and regulators
Auditors and regulators ask for what a good release process already produces: metrics and thresholds, test data and its provenance, results per release, approvals and production monitoring. For high-risk systems, Article 15 of the EU AI Act requires an appropriate level of accuracy, robustness and cybersecurity throughout the lifecycle, with accuracy levels and metrics declared in the instructions of use. Article 9 requires testing against predefined metrics and thresholds, Article 17 documented test and validation procedures before, during and after development, and Article 72 post-market monitoring.10
After the 2026 amendment, these obligations apply from December 2, 2027 for the high-risk uses in Annex III and from August 2, 2028 for AI in regulated products.12 Outside the EU, the NIST AI RMF sets the same expectations in its Measure function, and its Generative AI Profile of July 2024 adds actions for risks such as confabulation.115
The evaluation file for each release
- Golden set version, sampling method, label provenance and agreement between expert labelers.
- Metric definitions, thresholds and owner signatures, dated before testing.
- Results per segment with confidence intervals, and the cases that changed.
- Judge calibration results and the judge model version.
- Red-team results, and signed exceptions with their expiry dates.
- The online evaluation plan: what is sampled, the alert lines and who responds.
None of this needs a new department. It needs the golden set, the gates and the owners in place before the first prompt is tuned. That is how we run AI development: the evaluation suite is built first and handed over with the system, so your team can test the next change, and the next model, on its own.
Questions leaders ask
What is the difference between LLM evaluation and LLM observability?
Evaluation measures quality against a standard: it runs defined cases, scores the outputs and decides whether a release passes. Observability records what the system did in production, including traces, inputs, tool calls, latency and cost. The two meet in online evaluation, where a sample of production traces is scored with the same metrics and judges used before release, and every failure found there becomes a new test case.
What metrics are used to evaluate LLMs?
It depends on what the system does. Extraction and classification use field-level accuracy and format validity. RAG systems add retrieval recall, faithfulness of each claim to its source and correct abstention. Agents add task success by end state, trajectory checks on tools and arguments, and consistency across repeated runs. Every system also needs cost per completed task, latency at p95 and a check for harmful or leaked output.
How many test cases does an LLM evaluation dataset need?
Enough to detect the change you care about. At a 90% pass rate, 100 independent cases give a margin of error of about 6 points either way, and 400 cases about 3. Each segment with its own threshold needs its own count, related cases need more, and a strict gate such as an error rate under 1% needs at least 300 clean cases, by the statistical rule of three.
Is LLM-as-a-judge reliable?
It can be, once it is calibrated on your cases. A 2023 study found the strongest judge agreed with human experts on 85% of non-tied pairwise votes,2 while a 2025 study found even the best judges well behind human agreement on absolute scores and inclined to leniency.9 Use pass or fail verdicts with a rubric and a reference answer, measure agreement against expert labels, and keep people reviewing a sample.
How do you evaluate an AI agent?
Grade three things. The outcome: compare the records the agent left with the expected end state. The trajectory: check from the trace that it called the right tools with the right arguments, asked for approval where required and took no forbidden action. Consistency: run each task several times and report pass^k, the share of tasks it completes on every attempt, because one successful run proves little.8
How often should an LLM application be re-evaluated?
On every change that can alter an output, and continuously in production. Prompt edits, model versions, retrieval settings, tool definitions and policy files all run the golden set before merge. A sample of live traffic is scored every day, because hosted models, source documents and user behavior change after release.3 Review and refresh the golden set itself each quarter.
Sources
- The state of AI in 2025: Agents, innovation, and transformationMcKinsey & Company, November 5, 2025 (survey of 1,993 participants, fielded June 25 to July 29, 2025)
- Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaNeurIPS 2023 Datasets and Benchmarks Track, arXiv 2306.05685
- How is ChatGPT's behavior changing over time?arXiv 2307.09009, July 2023
- The 2026 AI Index ReportStanford Institute for Human-Centered Artificial Intelligence, 2026
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)National Institute of Standards and Technology, July 26, 2024
- Pervasive Label Errors in Test Sets Destabilize Machine Learning BenchmarksNeurIPS 2021 Datasets and Benchmarks Track, arXiv 2103.14749
- Adding Error Bars to Evals: A Statistical Approach to Language Model EvaluationsarXiv 2411.00640, November 2024
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsarXiv 2406.12045, June 2024
- Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-JudgesProceedings of the Fourth Workshop on Generation Evaluation and Metrics (GEM2), Association for Computational Linguistics, 2025
- Regulation (EU) 2024/1689, the Artificial Intelligence ActOfficial Journal of the European Union
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1National Institute of Standards and Technology, January 2023
- AI ActEuropean Commission, Shaping Europe's digital future, updated for the AI Omnibus, Regulation (EU) 2026/1744, in force since July 27, 2026
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on September 26, 2026. No client data appears in our insights.