· Prakash Natarajan · Reliability · 15 min read
AI Agent Evaluation Framework: What to Measure
Most agent evaluation guides stop at plan quality and tool correctness. Here's the fourth layer they skip: turning an eval score into a number your own customers can see.

An AI agent evaluation framework is the set of metrics, thresholds, and a labeled dataset that let you score whether an agent’s plan, tool calls, and final outcome are good enough to ship, run on every deploy the same way a test suite runs. Most write-ups on the topic stop at three layers: does the agent’s plan make sense, did it call the right tool with the right arguments, and did the task actually finish. That gets you a passing score nobody outside your own team ever sees. It says nothing about the number your own customers eventually ask for, how many of their tickets the agent actually closed, and what closing one cost. This piece covers the standard three layers, then adds the fourth one a B2B2C AI SaaS builder can’t skip: turning an internal eval score into a customer-facing metric, sliced correctly once more than one tenant runs behind the same agent.
What is an AI agent evaluation framework, and how is it different from evaluating a plain LLM call?
An AI agent evaluation framework scores a full run, not a single generation. Evaluating one LLM call means checking whether one prompt produced one good response. Evaluating an agent means checking a chain of decisions: which plan it picked, which tools it called along the way, whether each call used the right arguments, and whether the sequence actually solved the user’s problem, since a wrong step in the middle can still end in a correct-looking final answer that hides real damage underneath it.

DeepEval’s own guide names this gap directly and ships metrics for each part of the chain: PlanQualityMetric and PlanAdherenceMetric for the reasoning step, ToolCorrectnessMetric and ArgumentCorrectnessMetric for the tool-call step, and TaskCompletionMetric and StepEfficiencyMetric for the run as a whole, plus a general-purpose GEval and DAGMetric for anything custom. IBM’s own breakdown groups the same idea into evaluation methods rather than named classes: benchmark testing against a fixed dataset, human-in-the-loop review, A/B testing between agent versions, and “LLM-as-a-judge,” where a second model scores the first one’s output against a rubric. Both are describing the same three-layer shape from two different angles, and neither one is wrong. The distinction that actually matters for a production agent is that a single end-to-end pass or fail tells you almost nothing about which layer broke, so a real framework scores each layer separately and only rolls them up afterward.
What should you measure at the reasoning, action, and execution layers?
Score the plan, the tool calls, and the final outcome as three separate numbers, not one blended pass or fail, because a plan can be sound while the tool call underneath it fails, or a tool call can succeed while the overall task still doesn’t get solved. Blending them into one score tells you an agent is broken without telling you where.

| Layer | What it catches | Named metric example |
|---|---|---|
| Reasoning | A plan that skips a required step or picks the wrong strategy before any tool runs | Plan quality, plan adherence |
| Action | The right plan, executed with a wrong or hallucinated argument | Tool correctness, argument correctness |
| Execution | Every individual step correct, but the task still doesn’t resolve, or resolves in too many steps | Task completion rate, step efficiency |
Braintrust’s own framework adds a fourth category on top of these three that both IBM and DeepEval also cover in some form: safety, scored as prompt injection resilience, policy adherence rate, and bias detection, sitting alongside the other layers rather than folded into any one of them. IBM’s version of the same idea groups metrics slightly differently again: task-specific scores like success rate, error rate, cost, and latency; a function-calling bucket that flags wrong function names, missing parameters, hallucinated parameters, and unit mismatches; and a user-experience bucket covering CSAT and conversational flow. None of these three sources disagree on what to measure so much as on how to bucket it, which means the actual choice you have to make isn’t which taxonomy is correct, it’s which buckets map cleanly onto the tools your own agent actually calls, so you’re not scoring a category that never applies to your product.
How do you set a passing threshold instead of picking one on gut feel?
Set the threshold from the cost of being wrong in each direction, not from a round number that feels safe. A threshold that blocks too aggressively creates a support ticket every time a fine action gets rejected, and a threshold that’s too loose lets a bad tool call through into production, and every one of the write-ups we compared this piece against leaves that trade-off unaddressed.

Start from the specific tool, not the agent as a whole. A read-only lookup that occasionally returns a slightly stale answer costs you almost nothing when it’s wrong, so a lower threshold that lets more traffic through unblocked is the right call. A refund-approval tool call is the opposite case: a false pass costs you the refund amount, and a false block just costs a support ticket asking why the agent stalled, so the threshold should sit closer to strict, favoring the block over the pass whenever the two costs aren’t close. This is the same risk-tiering idea covered from the guardrail side in AI agent guardrails: a guardrail decides whether an action is allowed to run at all, and an eval threshold decides whether last week’s build is good enough to ship, but both draw the line from the same input, what a wrong answer actually costs on that specific tool. Once you’ve set an initial threshold this way, treat it as a starting point, not a fixed rule: track the false-positive and false-negative rate against real production outcomes for a few weeks and move the threshold toward whichever failure mode is actually showing up more often in your own traffic.
How do you build an eval pipeline that plugs into LangChain, LlamaIndex, or CrewAI?
Build the pipeline around three things: a labeled dataset pulled from real production traces, a judge that scores new runs against that dataset, and a gate that blocks a deploy when the score drops. Braintrust’s own product page claims native support across more than a dozen frameworks, including LangChain, LlamaIndex, Vercel AI SDK, OpenAI Agents SDK, and CrewAI, but stops at naming the integrations without describing what changes per framework, which is the exact gap this section fills.

The dataset comes first and matters more than the choice of judge model. Pull real traces from production, scrub anything tenant-identifying or personally identifying before it goes into a shared eval set, and label each one with the plan, tool calls, and outcome you’d have wanted, not the one you got. A hand-labeled set of thirty to fifty real cases, weighted toward the tool calls that carry real consequences, catches more regressions than a much larger set of synthetic ones, because synthetic cases tend to cluster around the easy paths an agent already handles well. Every framework in the list above exposes some form of hook after each step, a callback, a trace event, an instrumentation point, and that hook is where you attach the judge, whether it’s a second LLM call scoring the output against your rubric or a deterministic check against the labeled answer. Run the fast subset of this set on every pull request the same way you’d run a unit test suite, the pattern covered in more depth in AI agent testing, and reserve the full dataset, including the slower and more expensive judge calls, for a scheduled run or a pre-release gate rather than every single commit.
What breaks when one eval score has to cover more than one tenant?
A single aggregate eval score hides a tenant-specific regression, because a build that scores 92% across your whole traffic can still be failing badly for one customer whose usage pattern doesn’t look like the average. None of the frameworks we reviewed for this piece address multi-tenant scoring at all, since they’re built around one team evaluating one agent for itself, not a company running the same agent underneath hundreds of separate end customers.

The fix is to tag every trace with the tenant it belongs to at the point the eval runs, not after the fact, and slice the score by tenant before you roll it up into one number for the release notes. The same principle covered from the data layer in tenant isolation for AI agents applies here: a tenant ID that gets attached at query time but dropped from the eval pipeline means your aggregate pass rate can climb even while one customer’s agent quietly degrades, because their traffic is small enough to disappear inside the average. In practice this means storing the tenant ID alongside every labeled case in your dataset, running the judge per tenant slice as well as on the pooled set, and setting an alert on any single tenant’s score dropping by more than a fixed amount even when the overall number looks fine. A build that passes on aggregate and fails for your biggest customer is not a passing build, it’s a blind spot with a green checkmark on it.
How do you turn an eval score into a number your own customers can see?
An internal eval score and a customer-facing metric answer two different questions, and the biggest gap in every eval framework we looked at is that they only ever answer the first one. An eval score tells your own team whether the agent is good enough to ship. A customer-facing metric, deflection rate, cost per resolution, resolution rate, tells the AI SaaS builder’s own end customer whether the product is actually worth paying for, and none of DeepEval, IBM, or Braintrust’s write-ups connect the two.

The connection is direct once you look for it. A task completion score from the execution layer is close kin to resolution rate, the share of conversations an agent actually closes without a human stepping in, covered with its full formula in deflection rate versus resolution rate. A tool-correctness score that’s dropping for a specific tenant is an early warning for a coming drop in that same tenant’s resolution rate, days before the customer-facing number moves and someone opens a support ticket asking why. The practical move is to keep the two numbers next to each other rather than in separate systems: when a build’s internal eval score dips on a specific tool, check whether that tool’s tenant-sliced resolution rate is already trending down, and treat the internal score as the leading indicator it actually is. This is the exact bridge AiAgRe’s own dashboard is built around, tracing the same event that feeds your internal eval into a white-label view the AI SaaS builder can hand straight to its own customer, so the number that decides whether to ship and the number that proves the product works come from one pipeline instead of two disconnected ones.
Evaluating the agent and proving it to a customer are two different jobs
A framework that scores plan quality, tool correctness, and task completion tells you the agent is good enough to ship, and stopping there is where every general eval guide we compared this piece against stops too. It doesn’t tell a B2B2C AI SaaS builder’s own end customer anything, and that customer is the one deciding whether to keep paying for the product built on top of the agent. Closing that gap means carrying the same trace that feeds your internal eval score all the way through to a tenant-sliced, customer-facing number, deflection rate, cost per resolution, resolution rate, rather than treating the two as separate pipelines built by separate teams at separate times. AiAgRe traces every step an agent takes with tenant identity attached from the start, so the eval score your team ships against and the ROI number your customer sees come out of the same underlying data, not two systems you have to keep in sync by hand.
Frequently asked questions
What’s the difference between an AI agent evaluation framework and an LLM evaluation framework?
An LLM evaluation framework scores one generation against one prompt. An AI agent evaluation framework scores a full chain of decisions, the plan, every tool call along the way, and the final outcome, because a wrong step in the middle can still produce a correct-looking answer while doing real damage underneath it, something a single-generation score can’t catch.
What metrics should an AI agent evaluation framework track?
At minimum, one metric per layer: a reasoning-layer metric like plan quality or plan adherence, an action-layer metric like tool correctness or argument correctness, and an execution-layer metric like task completion rate. Safety metrics, prompt injection resilience and policy adherence, sit alongside these rather than replacing any of them, and a fourth, customer-facing layer, deflection rate, cost per resolution, resolution rate, closes the gap between an internal pass and a metric your own customer actually cares about.
How do you set a passing threshold for an AI agent eval?
Set it from what a wrong answer costs on that specific tool, not from a default percentage. A low-consequence tool, a read-only lookup, tolerates a looser threshold, while a tool with a real side effect, a refund, a database write, needs a stricter one, since a false pass there costs more than a false block. Track false positives and false negatives against real production outcomes for a few weeks and adjust the threshold toward whichever failure mode is actually showing up.
Does an AI agent eval framework work differently in a multi-tenant product?
Yes. An aggregate score across all tenants can look fine while one specific tenant’s agent is quietly failing, since their traffic is often small enough to disappear inside the average. Tag every trace with the tenant it belongs to at eval time, score each tenant’s slice separately as well as the pooled set, and alert on a single tenant’s score dropping even when the overall number holds steady.
How does an internal eval score connect to a metric like deflection rate?
A task-completion score from the execution layer moves in the same direction as resolution rate, the share of conversations an agent closes without a human. When a specific tool’s eval score drops for one tenant, that tenant’s resolution rate is often about to drop too, which makes the internal score a leading indicator for the customer-facing number rather than a separate metric tracked in a separate system.
Related reading: AI agent guardrails covers the threshold logic from the blocking side, AI agent testing covers building the pull-request-level test suite this piece’s pipeline plugs into, G-Eval for AI agents covers the LLM-as-judge scorer this pipeline’s evaluation step can run, tenant isolation for AI agents covers the data-layer side of the multi-tenant gap above, and deflection rate versus resolution rate covers the formula behind the customer-facing metric this piece connects an eval score to.