· Prakash Natarajan · Reliability · 16 min read

RAG Evaluation Metrics: What the Score Misses

Every RAG evaluation guide covers the same five metrics and stops at the CI log. Here's the concrete cost math, real thresholds, and the customer-facing gap none of them touch.

Every RAG evaluation guide covers the same five metrics and stops at the CI log. Here's the concrete cost math, real thresholds, and the customer-facing gap none of them touch.

RAG evaluation metrics split into two scorecards: retrieval metrics that check whether the right chunks came back, and generation metrics that check whether the answer actually used them honestly. The five metrics almost every framework converges on are context precision, context recall, faithfulness, answer relevancy, and answer correctness, each computed by a second model acting as a judge. None of that tells you what a real eval run costs to operate, where the pass and fail line actually sits, or how a customer of yours would ever see the result.

What Are RAG Evaluation Metrics, and Why Does RAG Need Two Scorecards Instead of One?

A RAG evaluation metric measures one specific way a retrieval-augmented generation pipeline can succeed or fail, and it needs two separate scorecards because the pipeline itself has two separate places to break.

two scorecards for a rag pipeline, one for retrieval and one for generation

The retrieval half pulls chunks out of a vector store or search index before the model ever sees the question answered. The generation half takes those chunks plus the user’s question and writes a reply. A single end to end score, “did the final answer look right,” can’t tell you which half actually failed. An agent that answers correctly because it happened to already know the fact from training data looks identical, on a blended score, to one that pulled the right chunk and used it properly, until the day the fact changes and only the second agent updates its answer. Splitting the score in two forces you to check retrieval and generation independently, which is the only way to know whether a bad answer came from a bad search or a model that ignored good context it was handed.

Every serious RAG evaluation guide, from the open source RAGAS project to newer entrants like DeepEval and Patronus AI, organizes its metrics around this same retrieval versus generation split, and that much is table stakes. Where they stop is at the CI log: none of the widely cited guides puts a defensible number on what a nightly eval run actually costs to execute, spells out where the pass or fail line should sit instead of a token example buried in code, or explains what happens to the score once it leaves your own dashboard. Those three gaps are what the rest of this piece fills.

How Do You Score Retrieval Quality?

Retrieval quality is scored by checking whether the chunks a query pulled back are the ones a correct answer would actually need, using context precision and context recall as the two standard metrics.

retrieval quality metrics for a rag pipeline shown as a ranked chunk list

Context precision asks what share of the retrieved chunks were actually relevant, and specifically whether the relevant ones landed near the top of the ranking rather than buried at position eight of ten. A retriever that returns three good chunks and seven irrelevant ones scores low on precision even if one of those three good chunks would have been enough, because the model still has to wade through the noise, and noise is exactly what causes a generation model to hedge or hallucinate. Context recall flips the question around: of everything a correct answer would need, how much did the retriever actually surface. A retriever that returns one perfectly relevant chunk out of three required ones scores well on precision and badly on recall, which is the classic failure mode of an overly narrow embedding search or a chunk size set too small to capture a full answer. Some teams also track simpler ground truth metrics borrowed from classic information retrieval, precision at k and recall at k, which measure the same idea against a fixed-size result list rather than an LLM’s judgment of relevance. All of these need either a labeled set of correct chunks per question, which is expensive to build by hand, or an LLM judge that reads the retrieved chunks and estimates relevance itself, which is cheaper but only as reliable as the judge model you pick.

How Do You Score Generation Quality?

Generation quality is scored by checking whether the model’s answer stays grounded in the retrieved context and actually addresses the question, using faithfulness and answer relevancy as the two standard metrics, with answer correctness as a third when you have a reference answer to compare against.

generation quality metrics for a rag pipeline shown as a grounded answer next to source chunks

Faithfulness, sometimes called groundedness or the inverse of a hallucination rate, breaks the generated answer into individual factual claims and checks each one against the retrieved context: a claim the context supports counts as grounded, a claim the context contradicts or never mentions counts against the score. This is the metric that catches the specific RAG failure mode everyone worries about, a model that retrieves the right chunks and then writes something plausible-sounding that those chunks never actually said. Answer relevancy works from the opposite direction, checking whether the answer actually addresses what was asked rather than drifting into a technically-grounded but off-topic reply, the kind of failure where every sentence is true but none of them answer the question. Answer correctness is the strictest of the three because it needs a reference answer to compare against, scoring how close the generated reply comes to that reference in both meaning and factual content, which makes it the right metric for a fixed regression suite and the wrong one for live production traffic where there’s no pre-written correct answer to check against.

What Do RAGAS, DeepEval, and TruLens Actually Compute Differently?

RAGAS, DeepEval, and TruLens compute the same core metrics under different names and different judge-model defaults, so the real difference between them is integration shape and how each one lets you swap in your own judge model rather than the underlying math.

three open source rag evaluation frameworks compared as parallel pipelines

RAGAS is the framework most of the industry vocabulary traces back to, the original source of “context precision” and “context recall” as named metrics, and it ships as a lightweight Python library you point at a dataframe of questions, retrieved contexts, and generated answers. DeepEval, built by Confident AI, wraps a similar metric set in a testing-framework shape closer to pytest, which makes it the more natural fit for a team that wants RAG checks running inside an existing CI pipeline rather than a standalone notebook, and it adds first-class support for a custom G-Eval style metric when the five standard scores don’t cover something specific to your product. TruLens leans harder into tracing, instrumenting each step of the pipeline so a failing score comes with the actual retrieval and generation trace attached rather than just a number, which shortens the loop from “the score dropped” to “here’s the exact chunk that caused it.” None of the three requires you to pick a specific judge model; all three default to whatever OpenAI or Anthropic model you configure, and the metric definitions stay the same regardless of which judge you point them at, only the judge’s own reliability changes. The practical choice between them usually comes down to where your team already lives: a pytest-based test suite favors DeepEval, a notebook-first research workflow favors RAGAS, and a team that needs full trace visibility into a production incident favors TruLens.

What Score Actually Means Your RAG Pipeline Is Failing?

A faithfulness score below roughly 0.75 or a context recall score below roughly 0.7 is a reasonable starting line for calling a RAG pipeline broken enough to block a release, though every team should tune these against its own tolerance for risk rather than treating them as a fixed industry standard.

a faithfulness score gauge with a shipped threshold and a blocked threshold marked

Most of the public RAG evaluation guides show a threshold in a code sample, commonly 0.7, without ever explaining why that number and not another one. Here’s the reasoning worth applying instead of copying the example verbatim. Faithfulness scores the share of claims in an answer that trace back to the retrieved context, so a score of 1.0 means every claim is grounded and a score of 0.75 means roughly one claim in four is not, which for a customer support agent quoting a return policy or a refund amount is already too much invented content to ship. Treat anything under 0.9 as worth triaging by hand and anything under 0.75 as a release blocker, not a warning. Context recall matters differently because it caps what faithfulness can even achieve: if the retriever only surfaced 70 percent of what a correct answer needed, no amount of careful generation can produce a fully grounded reply, so a recall floor around 0.7 is really a retrieval-layer gate that should trigger before you spend a token on generation at all. The one number that resists a universal threshold is answer relevancy, because a deliberately terse or a deliberately thorough house style will both score differently on the same underlying quality, so calibrate it against your own answers rather than a borrowed cutoff.

What Does Checking That Score Actually Cost at Production Scale?

Running the five standard RAG evaluation metrics with an LLM judge costs roughly a tenth of a cent per test case per metric on a fast, cheap judge model, which sounds trivial until you multiply it by every metric, every case, and every night the suite runs.

a cost meter for llm as judge rag evaluation runs at increasing scale

Work the actual math instead of trusting the “it’s cheap” hand-wave every guide leans on. Claude Haiku 4.5, a fast, inexpensive model well suited to judge work, currently prices at one dollar per million input tokens and five dollars per million output tokens. A single faithfulness check, feeding the judge the retrieved context, the generated answer, and a short rubric, runs around 800 input tokens and 150 output tokens once the judge’s reasoning is included, which comes out to roughly $0.00155 per check. Multiply that by all five standard metrics and a 200-item regression suite run every night, and the nightly bill lands around $1.55, or about $47 a month, which really is close to nothing. The number that actually matters is production sampling, not the regression suite. A support agent handling 50,000 conversations a day, evaluating just 5 percent of them live in production across the same five metrics, runs 12,500 judge calls a day, roughly $19 a day or $580 a month. That’s still a rounding error against the cost of the underlying agent traffic, but it’s not the “basically free” the docs imply, and a team that skips this math before scaling sample rate to catch more incidents can be surprised by the line item three months in. Sample at a rate that matches your actual incident-detection need, not the highest rate you can afford, and revisit it once you’ve seen a quarter of real drift data.

How Do These Scores Reach the Customer Who’s Actually Using the Agent?

These scores reach the customer only if someone builds the pipeline connecting an eval result to a metric that customer can actually see, and none of the RAG evaluation guides on the first page of results even raises the question.

a rag evaluation score rolling up into a customer facing dashboard tile

Every framework covered above, RAGAS, DeepEval, TruLens, Patronus, stops at the moment a score lands in a CI log or an internal dashboard your own engineering team checks. That’s the right scope for those tools, and it’s also exactly the gap for a B2B2C AI SaaS team: if you’re shipping a support or sales agent built on RAG, it’s not your engineers who need to trust the retrieval layer, it’s the end customer using your product to run their own business. A faithfulness score dropping from 0.94 to 0.81 on one tenant’s document set is a leading indicator that tenant’s resolution rate is about to fall, days before a support ticket asks why the agent started making things up, covered in more detail in deflection rate, explained. Most teams keep these as two separate systems, an eval dashboard the engineering team checks and a metrics dashboard the account team checks, and lose the leading-indicator relationship between them in the process, the same gap covered from the testing side in the AI agent evaluation framework. AiAgRe closes it by scoping every traced event to org and end-customer from ingestion, so the same trace that feeds a faithfulness or context recall check also rolls up into the deflection percentage, cost saved, and resolution rate a white-label dashboard component shows the AI SaaS builder’s own customer, the pattern covered in full in customer-facing analytics. Multi-tenant scoping matters here specifically because a RAG pipeline shared across tenants can score well on aggregate while one tenant’s document set quietly drifts underneath the average, a failure mode covered in tenant isolation for AI agents.

Ship the score before the tenant finds the gap

Every RAG evaluation write-up this piece was checked against gets the core metrics right: split retrieval from generation, run faithfulness and context recall as the two load-bearing scores, and pick RAGAS, DeepEval, or TruLens based on where your team already works. None of them tell you the number that actually decides whether a pipeline ships, what running that check costs once you move past a toy regression suite, or what happens to the score after your own team stops looking at it. Set a faithfulness floor you’d actually block a release on, budget the judge-model cost against your real production sample rate instead of a regression suite alone, and connect the score to whatever your end customer already sees, because a RAG pipeline that scores 0.95 in your CI log and drops to 0.7 on one tenant’s real documents has not actually been evaluated where it matters. AiAgRe traces every retrieval and generation step with tenant identity attached from the first event, so the same data behind your RAGAS or DeepEval run rolls straight into the customer-facing reliability number your dashboard already shows, instead of asking you to wire the two together by hand.

Frequently asked questions

What’s the difference between context precision and context recall?

Context precision measures what share of the chunks a retriever returned were actually relevant, penalizing noise and poor ranking. Context recall measures what share of the chunks a correct answer needed were actually returned, penalizing a retriever that missed something important. A pipeline can score high on one and low on the other, which is exactly why both get tracked separately rather than blended into one number.

Do you need all five RAG evaluation metrics, or can you pick a subset?

Pick a subset if the extra metrics are not answering a question you actually have. Faithfulness and context recall are the two load-bearing scores for most teams because they catch hallucination and missing information respectively. Answer correctness only matters if you have reference answers to compare against, and context precision matters most when retrieval noise is a known problem rather than a hypothetical one.

How often should you run RAG evaluation in production?

Run the full regression suite on every deploy, since it’s cheap and catches a broken build before it ships. Run live production sampling at a rate tied to how much undetected drift you can tolerate, not the highest rate you can afford, since the cost scales linearly with sample rate and conversation volume.

Can you evaluate RAG without a reference or ground-truth answer?

Yes, for four of the five standard metrics. Context precision, context recall, faithfulness, and answer relevancy can all run reference-free, using an LLM judge to assess relevance and groundedness directly against the retrieved context rather than a pre-written correct answer. Only answer correctness genuinely needs a reference to compare against.

Does agentic RAG need different evaluation metrics than plain RAG?

It needs the same retrieval and generation metrics plus additional ones, not a replacement set. An agent that loops through multiple retrieval and tool calls before answering also needs task completion and argument correctness checked at each step, since a wrong tool call three steps into a trajectory can still produce a fluent, well-grounded final answer that the standard five metrics would score as a pass, a gap covered in more detail in RAG vs agentic RAG.

Related reading: RAG vs agentic RAG covers the added failure modes once retrieval feeds a looping agent instead of a single reply, the AI agent evaluation framework covers the reasoning, action, and execution layers a RAG check plugs into, tenant isolation for AI agents covers the data-layer side of the multi-tenant scoring gap above, and deflection rate, explained covers the customer-facing metric a faithfulness drop is an early warning for.

Back to Blog