· Prakash Natarajan · Reliability · 15 min read
How to Build a Golden Dataset for an AI Agent
Most golden dataset guides assume one team scoring one flat input-output pair. An AI agent needs the full trajectory, tagged per tenant, and tied to the number your own customers see.

A golden dataset is a fixed, versioned set of labeled cases that you replay against every new agent build before it ships, the same way a test suite runs before every deploy. For a plain LLM call, one item is a prompt and the reply you’d want back. For an agent, that’s not enough, because a wrong tool call two steps into a run can still end in a reply that reads correctly while the actual outcome is wrong, so a real item has to capture the whole trajectory: the plan, every tool call along the way, and the final result. Most write-ups on the topic are written for one team scoring one generic LLM call and stop there. This piece covers the standard build-and-maintain lifecycle, then adds the two things a B2B2C AI SaaS builder can’t skip: scoping the dataset so one tenant’s regression doesn’t hide inside everyone else’s average, and connecting the score to the number your own customer actually sees.
What is a golden dataset for an AI agent?
A golden dataset for an AI agent is a hand-curated, versioned set of test cases, each one built from a real or realistic user request plus the trajectory and outcome a correct agent run should produce, that you run every new build against before it reaches production.

The term gets used loosely outside of AI too, and that’s worth clearing up before going further. In the business intelligence world, a golden dataset usually means a single, trusted source of truth a company uses for reporting, the one number everyone agrees is correct rather than five conflicting exports. In LLM evaluation, the meaning narrows: it’s a fixed comparison set, not a live source of truth, and its job is regression detection, not reporting. DeepEval’s own documentation calls a golden a “precursor to a test case”, something that only needs an input or scenario to initialize and stores the expected result as expected_output or expected_outcome. It also splits the idea in two: a plain Golden for single-turn cases, and a ConversationalGolden for multi-turn ones, which adds a scenario, an expected_outcome, and optional persona and context fields. That split matters for an agent, since most real agent runs are conversational and multi-step rather than one prompt and one reply, and a dataset built only around single-turn goldens will quietly miss the failure modes that only show up three or four turns in.
What actually belongs in a golden dataset item, beyond input and output?
A golden dataset item for an agent needs four things, not two: the input or scenario, the expected plan or reasoning path, the expected tool calls with their arguments, and the expected final outcome, because scoring only the final reply lets a wrong step in the middle hide behind a final answer that happens to look right.

DeepEval’s Golden and ConversationalGolden schemas are the closest thing to a standard here, and they’re a real improvement over a flat prompt-and-reply pair, but both still center on what the model said back, not on the path it took to get there. An agent item needs to go further: tag the plan the agent should have chosen when more than one strategy is plausible, record the exact tool name and arguments a correct call should carry, and note the outcome state the run should end in, not just the sentence describing it. A refund tool called with the wrong order ID and a refund tool never called at all can produce nearly identical final replies to a user, “I’ve processed that for you”, and only a trajectory-aware item catches the difference between the two. The same gap shows up in a booking agent that reschedules the wrong appointment while still replying with a confirmation message that reads exactly like a correct one, or a lookup agent that calls a cached, stale record instead of the live one and never mentions the difference in its final answer. This is the same three-layer shape covered in more depth in the AI agent evaluation framework: a golden dataset is what supplies the labeled cases that framework scores against, and an item that’s missing the plan or the tool-call layer leaves two of those three layers with nothing real to check.
How do you build your first golden dataset?
Build the first version from real production traces, not synthetic examples, because synthetic cases tend to cluster around the paths an agent already handles well and miss the edge cases that actually cause tickets.

Langfuse’s own workflow for this is concrete enough to copy directly: every trace and observation carries an “Add to dataset” action in the UI, and for bulk conversion you filter observations by score, feedback, or tag, select the ones worth keeping, and run an Actions to add-to-dataset step that maps the trace’s fields onto a dataset item. The same thing happens programmatically with one SDK call, create_dataset_item(), passing a dataset name, the input, the expected output, and the source trace ID so the item stays traceable back to where it came from. Once production traffic gets thin for a specific tool or edge case, deepeval’s Synthesizer fills the gap by generating single-turn goldens from documents or existing context, and its ConversationSimulator generates multi-turn conversation turns for cases you can’t yet pull from real usage, though both should stay the minority of the set rather than the majority of it. Weight whatever mix you end up with toward the tool calls that carry real consequences rather than spreading items evenly across every tool the agent has: a handful of well-labeled cases around a refund or a database write catches more regressions than the same number of cases spread evenly across a dozen low-stakes lookups, because that’s where a wrong call actually costs something.
Size the dataset by what you’re using it for rather than reaching for one number that covers every case:
| Purpose | Rough size |
|---|---|
| Exploring one specific issue | About 10 items |
| Testing a model swap or capability change | About 10 complex, representative cases |
| A pull-request gate that has to run fast | Tens to low hundreds of items |
| A full CI check on a larger change | 100 to 1,000 items |
A pull-request gate that takes ten minutes to run stops getting run, so keep that subset small and fast, and reserve the larger set for a scheduled or pre-release pass, the same split covered from the pipeline side in AI agent testing.
How do you scope a golden dataset across tenants?
Tag every golden dataset item and every score with the tenant or vertical it belongs to at write time, not after the fact, because a single pooled dataset can pass on aggregate while quietly missing a failure mode that only shows up for one customer’s specific tool configuration or business language.

None of the write-ups this piece was checked against, not Langfuse’s, not deepeval’s, not the generic BI definitions, address multi-tenant scoping at all, because they’re built around one team evaluating one agent for itself rather than one agent instance running underneath hundreds of separate end customers with different tool integrations and different vocabularies. The practical fix is to keep two layers in the same dataset rather than one flat pool: a shared core subset that every tenant’s build has to pass, covering the trajectory shapes that apply everywhere, and a per-tenant or per-vertical subset built from that tenant’s own production traces, covering the tool configurations and phrasing that don’t generalize across customers. This is the same principle covered from the data-isolation side in tenant isolation for AI agents: a tenant ID attached to a trace but dropped before it reaches the eval pipeline means a regression specific to your biggest customer can sit underneath a passing aggregate score for weeks before anyone notices, because their traffic is a small enough slice of the pooled set to disappear inside the average.
How do you keep a golden dataset from going stale?
A golden dataset goes stale the moment the agent’s real traffic drifts away from what the dataset was built to check, so catching it means watching for a gap between strong scores on the dataset and weakening signals from actual production, not waiting for the dataset to obviously break.

Langfuse names that exact symptom directly: drift shows up as “divergence between strong experiment scores and weakening production signals such as user feedback or online evaluator scores”, which is a quieter failure than a dataset that just stops running. The fix is ongoing rather than one-time: keep converting new production failures into dataset items the same way the first version got built, record when each item was added so an aging cluster is visible at a glance, and retire items that no longer reflect how the agent is actually used, since a dataset that only ever grows accumulates dead weight that slows every run without adding coverage. Deduplication matters here too, because twenty near-identical phrasings of the same underlying question overweight that one case in every average score the dataset produces, so check whether an intent is already covered before adding another near-copy of it. Langfuse versions every addition, update, deletion, or archival automatically, timestamped with item-level diffs, which is what makes it possible to pin a dataset version when comparing scores across two time periods instead of accidentally comparing a build against a dataset that quietly changed underneath it. On the CI side, Langfuse ships a GitHub Action, langfuse/experiment-action, that loads a dataset, runs the experiment, posts the scores as a pull-request comment, and fails the check on a regression, wiring the whole maintenance loop into the same place engineers already look.
How does a golden dataset connect to what your customers actually see?
A golden dataset score and a customer-facing metric answer two different questions, and the gap between them is the biggest thing every general eval write-up leaves alone: a dataset score tells your own team whether a build is good enough to ship, while a metric like deflection rate or resolution rate tells the AI SaaS builder’s own end customer whether the product they’re paying for is actually working.

The connection is direct once the same trace ID runs through both. A tool-call item failing for a specific tenant’s golden dataset subset is an early warning for that same tenant’s resolution rate, the share of conversations the agent closes without a human stepping in, covered with its full formula in deflection rate, explained, days before the drop shows up in a support ticket asking why. Most teams keep the two in separate systems, an eval dashboard the engineering team checks and a metrics dashboard the account team checks, and lose the leading-indicator relationship between them in the process. AiAgRe is built around closing that gap specifically: its Node SDK scopes every traced event to org and end-customer from ingestion, so the same event that feeds a golden dataset comparison also rolls up into the deflection percentage, cost saved, and resolution rate a white-label dashboard component shows the AI SaaS builder’s own customer, the pattern covered in more detail in customer-facing analytics. The dataset score and the ROI number aren’t two separate pipelines built by two separate teams at two separate times, they’re two views of the same trace.
Ship the dataset before the agent needs saving
A golden dataset earns its keep the first time it catches a regression before a customer does, and every write-up this piece was checked against gets the standard lifecycle right: pull real traces into items, size the set to what it’s being used for, keep it versioned, and wire it into CI so a bad build never ships quietly. None of them account for an agent’s full trajectory instead of a flat reply, for more than one tenant sharing the same evaluation pipeline, or for the fact that a dataset score only matters to your own team unless it’s connected to a number your customer can see. Build the dataset with those three gaps closed from the start, and it stops being a QA artifact only your engineers look at and becomes the thing standing between a bad deploy and the AI SaaS builder’s own customer finding out first. AiAgRe traces every step of an agent’s run with tenant identity attached from the first event, so the same data that feeds your golden dataset comparisons rolls straight into the customer-facing metrics your dashboard already shows, instead of asking you to stitch the two together by hand.
Frequently asked questions
What is a golden dataset in AI, in one sentence?
A golden dataset is a fixed, versioned set of labeled test cases, each one pairing a real or realistic input with the trajectory and outcome a correct run should produce, that gets replayed against every new build to catch a regression before it ships.
How many items does a golden dataset need?
It depends on what you’re using it for. About ten items is enough to explore one specific issue or test a model swap, a pull-request gate that has to run fast should stay in the tens to low hundreds so the check doesn’t slow every commit, and a full CI check on a larger change can reasonably run against a hundred to a thousand items.
What’s the difference between a golden dataset and a benchmark?
A benchmark is usually a public, shared test set used to compare different models or systems against each other on a general capability. A golden dataset is private and specific to your own agent and your own product, built from your real traffic and your real tool integrations, and its job is catching a regression in your build, not ranking your model against someone else’s.
Do you need a separate golden dataset per tenant?
Not entirely separate, but tagged and scored separately. Keep a shared core subset every tenant’s build has to pass, covering the trajectory shapes that hold across your whole product, and a per-tenant subset built from that tenant’s own traces, covering the tool configurations and phrasing that don’t generalize, since a single pooled score can look fine on aggregate while one customer’s agent quietly fails underneath it.
How do you know when a golden dataset has gone stale?
Watch for a gap between the dataset’s scores and real production signals, strong experiment results next to weakening user feedback or online evaluator scores, rather than waiting for the dataset to obviously break. Keep converting new production failures into items, record when each item was added, and retire the ones that no longer reflect how the agent is actually used.
Related reading: the AI agent evaluation framework covers the reasoning, action, and execution layers a golden dataset supplies the cases for, AI agent testing covers the pull-request-level pipeline this piece’s CI gate plugs into, tenant isolation for AI agents covers the data-layer side of the multi-tenant scoping gap above, and deflection rate, explained covers the formula behind the customer-facing metric this piece connects a dataset score to.