· Prakash Natarajan · Reliability · 18 min read

Context Rot in AI Agents: How to Catch It Early

Context rot degrades an AI agent's accuracy as a conversation grows, even with room left in the context window. What causes it, when it starts, and how to test for it.

Context rot degrades an AI agent's accuracy as a conversation grows, even with room left in the context window. What causes it, when it starts, and how to test for it.

Context rot is the drop in a model’s accuracy as the input it has to read gets longer, even when every fact it needs is still sitting somewhere inside that input. In a production AI agent, it shows up as a conversation that starts sharp and drifts: the agent forgets an instruction from three tool calls ago, calls the right function with a slightly wrong argument, or narrates a made-up outcome in the same confident tone it would use for a real one. It isn’t memory loss, and it isn’t a context window running out of room. It’s attention quietly thinning out over a window that still has plenty of space left, and it gets worse the longer a support copilot or sales assistant keeps a customer talking.

What is context rot in an AI agent?

Context rot is what happens when a language model has to search through a long input to find the piece it actually needs, and the search itself becomes the point of failure, not the absence of the fact. The information is technically present, but the model just pays less attention to it the further it sits from the start or the end of what it’s reading.

context rot definition, a highlighted fact buried in a long context window

Researchers usually describe this through positional bias: a model recalls information placed at the very beginning or the very end of its context far more reliably than the same information placed somewhere in the middle. Chroma’s research team ran this test directly, evaluating eighteen models including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 across several controlled tasks designed to isolate input length as the only variable. One of their clearest results came from a long conversational memory benchmark: a short, focused prompt of roughly three hundred tokens dramatically outperformed the same underlying task wrapped in a full transcript running past a hundred thousand tokens, even though the full transcript contained every fact the focused prompt did. The distance between “the model has the answer somewhere in its input” and “the model can actually find and use that answer” is exactly what context rot measures.

For an AI agent, this distinction matters more than it does for a one-off question to a chatbot. A support copilot or sales assistant doesn’t get to ask one clean question and stop. It runs a growing multi-turn conversation, often stitched together with tool calls, retrieved documents, and earlier turns the customer already forgot they said. Every one of those pieces adds tokens to the context the model has to search on the very next turn, and none of them get any easier to find as the pile grows.

At what point does a conversation start to rot?

There’s no single fixed token count where a model suddenly breaks. Degradation is gradual and it starts earlier than most teams assume, well before a model gets anywhere near its stated context limit.

an accuracy curve declining as input length grows in a long ai agent conversation

The clearest early evidence for this comes from Stanford’s “lost in the middle” research, which tested how accurately a model could retrieve one fact from a set of retrieved documents depending on where that fact sat in the input. With around twenty documents totaling roughly four thousand tokens, accuracy for a fact placed at the very start of the input landed around 70 to 75 percent. The same fact placed in the middle of that same twenty-document set dropped accuracy to roughly 55 to 60 percent, a fifteen to twenty percentage point swing driven entirely by position, not by anything wrong with the fact itself. Four thousand tokens is not a large input by current standards. It’s closer to a support conversation running eight or ten turns deep with a couple of retrieved help articles attached, which is an entirely ordinary shape for a customer support session by the time it’s actually resolved something.

Chroma’s own testing found a related pattern with repeated content: as input length grew past roughly five hundred to seven hundred and fifty words in one of their synthetic tasks, some models started generating output that had drifted from the actual input entirely, effectively guessing rather than reading. Different model families fail differently once this starts. Chroma found Claude models tend to abstain or hedge under uncertainty once accuracy starts slipping, producing a cautious non-answer rather than a wrong one, while GPT models were more likely to answer confidently anyway, which is the more dangerous failure mode for a customer-facing agent because a wrong, confident answer is much harder to catch than a model admitting it isn’t sure.

The practical takeaway isn’t a magic token number to design around. It’s that “rot starts somewhere past ten thousand tokens” is a far more useful working assumption for a multi-turn agent than “rot starts near the context limit”, because the second assumption leaves a huge stretch of ordinary, everyday conversations exposed without anyone realizing it.

How does context rot actually show up in a live agent?

Context rot rarely announces itself as an error. It shows up as a fluent, plausible-sounding reply that quietly did the wrong thing, which is exactly why it survives a quick read of the transcript and only gets caught by someone who checks the actual trace.

a live agent transcript destabilizing partway through a long conversation

The most common pattern is an agent re-asking for information the customer already gave several turns earlier, because the instruction to “check what the user already told you” got buried in the middle of a growing transcript and lost priority against more recent tokens. A close second is a tool call that fires with a slightly wrong argument, one pulled from an earlier, now-stale part of the conversation instead of the customer’s latest correction, so the call technically succeeds while quietly solving the wrong problem. The third pattern is the most expensive one: a ghost action, where the agent tells the customer something happened (a refund processed, a ticket escalated) without ever actually calling the tool that would have made it true, because the system instruction defining what “done” requires got diluted by everything piled on top of it by that point in the conversation.

None of these failures look unusual in isolation. A customer reading the reply sees a fluent, on-topic answer. An engineer skimming the final message sees the same thing. The only way to catch any of the three is to read the trace behind the reply and check what the model actually did against what it claimed, which is the same discipline covered in more depth in AI agent testing and in what a well-instrumented trace should capture in AI agent observability.

How do you catch context rot before a customer does?

You catch it by testing the same conversation at increasing lengths and watching where accuracy starts to drop, rather than trying to detect it after the fact from live traffic using statistical drift methods built for a data science team.

an evaluation grid testing pass and fail results across increasing conversation length

Some teams reach for embedding-drift classifiers or distribution-distance metrics to spot context rot in production, and those methods work, but they need a dedicated ML background to set up correctly and they tell you rot is happening somewhere without telling you exactly where in a specific conversation it broke. A smaller team building a support copilot or sales assistant doesn’t need that machinery to get real signal. Take five or six of your agent’s real capabilities, and for each one, build a short conversation that plants one specific fact early on: a customer’s plan tier, a case number, a constraint they stated once. Then extend that same conversation with realistic filler turns, retrieved documents, and tool call results, testing whether the agent can still correctly use that early fact at three or four increasing lengths, say two thousand, eight thousand, and twenty thousand tokens. Record exactly where the correct answer starts slipping.

This test earns its keep because it’s cheap to build once and reusable every time you change the underlying model or rewrite the agent’s instructions, which is exactly when a context rot regression is most likely to appear unannounced. Run it the same way you’d run any other eval in your suite: a code-based check wherever the correct outcome has an exact, checkable shape (the right plan tier gets referenced, the right case number gets used in a tool call), and a calibrated model-as-judge only where the answer is more open-ended. Pair the result with the length at which it broke, because “this fails around twelve thousand tokens” is a finding you can act on, while “this sometimes fails” is not.

How do you fix context rot without buying a new database?

You fix it by keeping what the model actually has to search through small and current, not by expanding the context window further or bolting on a new storage layer as the first move.

a long conversation history being compressed into a short rolling summary

The most direct fix is rolling summarization: instead of forwarding the full, growing transcript on every turn, periodically compress everything older than the last few exchanges into a short, dense summary, and forward that summary plus the recent turns instead of the entire history. LangChain, LlamaIndex, and CrewAI all support this pattern natively, whether through a summarization memory component, a checkpoint step in a graph-based agent, or a custom truncation hook before the next model call. The mechanics differ by framework, but the goal is identical: the model should never have to search a hundred thousand tokens to find a fact that could have been carried forward in three hundred.

Placement matters as much as compression. Given the positional bias behind context rot, a fact that absolutely cannot be lost, an account ID, a hard constraint, a policy the agent must never violate, belongs restated near the end of the prompt on every turn, right next to the most recent user message, not buried once near the top and left to compete with everything added after it. This is a small, nearly free change and it directly targets the exact mechanism the Stanford research measured.

The instinct to solve this by adding a bigger, smarter memory store is understandable but usually premature. A persistent memory layer solves a different problem, remembering facts across separate sessions, which AI agent memory covers in more depth. Context rot is a within-conversation problem: the fix is trimming and restating what the model reads on the very next call, not storing more for later. Reach for external memory once you’ve confirmed rolling summarization and smarter placement aren’t enough, not as the first move.

What does context rot cost you in deflection rate and resolution rate?

Context rot doesn’t just produce an occasional bad reply. It quietly corrupts the exact numbers you report to your own customers, because a deflection rate or resolution rate is only as trustworthy as the conversations feeding it, and long conversations are precisely where context rot concentrates.

a customer facing metrics dashboard showing deflection, cost, and resolution figures moving the wrong direction

Walk the chain through. A ghost action, where the agent claims a refund or escalation happened without the tool call that would have made it true, gets logged as resolved by whatever classifier trusts the agent’s own closing message, inflating your resolution count with conversations that were never actually finished. A tool call fired with a stale argument pulled from an earlier, now-outdated part of the conversation still returns a plausible-looking result, so it counts as a successful deflection even though it solved a slightly different problem than the one the customer actually had. And the interrogation-loop pattern, where an agent re-asks for information it already has, doesn’t fail outright, but it burns extra turns and extra tokens on every occurrence, quietly raising your true cost per resolution the way described in more detail in deflection rate, explained.

The fix connects directly to the testing work above. If your product marks a conversation “resolved” based on a closing-intent signal or an explicit tool call, tie at least one of your length-based eval tests to that exact classification logic, not just to whether the final reply reads well. A test that catches a ghost action at fifteen thousand tokens is the same test that keeps your reported resolution rate honest at that length, and it’s a cheap thing to add once the shorter-conversation version of the same test already exists in your suite. AiAgRe’s tracing sits underneath exactly this layer: because every model call and tool call gets captured at the trace level and tied to the same deflection, resolution, and cost-per-resolution definitions your white-label dashboard renders, a length-based eval that catches rot in testing is reading from the same evidence the number on your customer’s screen eventually shows.

Test one long conversation this week

Pick your agent’s single longest, most common real conversation type and rebuild it as a test: plant one specific fact early, extend the conversation with realistic filler out to two or three increasing lengths, and check whether the agent can still use that early fact correctly at each length. Note the exact point where it starts to slip rather than settling for “it’s usually fine”. Then apply the cheapest fix first: compress everything older than the last few turns into a short rolling summary, and restate anything that absolutely can’t be lost near the end of the prompt, closest to the model’s most recent input, rather than reaching for a bigger context window or a new memory system as the first move.

None of this requires a data science team or a drift-detection pipeline. It requires one honestly long test conversation, a habit of checking where accuracy actually breaks instead of assuming the context window is the only limit that matters, and a fix that trims what the model reads rather than expanding it further. If you’re building the kind of AI product where your own customers will eventually look at the deflection rate and resolution numbers your agent produces, AiAgRe ties that same trace data into a white-label dashboard, so the conversation length that broke your eval test is visible in the same place the number your customer trusts comes from.

Frequently asked questions

What is context rot?

Context rot is the drop in a language model’s accuracy as its input gets longer, even when the fact it needs is technically present somewhere in that input. It’s driven by positional bias: models recall information near the start or end of their context far more reliably than information buried in the middle, so performance degrades well before the context window actually runs out of room.

Is context rot the same thing as running out of context window?

No. Running out of context window means the input no longer fits and has to be truncated or rejected outright. Context rot happens while there’s still plenty of room left. The model technically has access to every fact but searches through the longer input less reliably, which is a retrieval and attention problem, not a capacity problem.

How many tokens does it take before context rot starts?

There’s no single fixed threshold, but degradation starts earlier than most teams expect. Stanford’s “lost in the middle” research measured a fifteen to twenty percentage point accuracy drop for a fact placed in the middle of roughly four thousand tokens of input, which is closer to a single support conversation than most teams assume. Treating ten thousand tokens as a realistic point where rot risk becomes real, rather than the model’s full context limit, is a safer working assumption for a multi-turn agent.

Does context rot affect every AI model the same way?

No. Chroma’s research found that different model families fail differently once accuracy starts slipping under long context. Claude models were more likely to abstain or hedge once uncertain, producing a cautious non-answer, while GPT models tended to answer confidently anyway, which is the harder failure mode to catch in a customer-facing agent because a wrong, confident answer looks the same as a correct one until someone checks the trace.

Can retrieval-augmented generation fix context rot?

Retrieval-augmented generation reduces how much irrelevant content ends up in the context window in the first place, which helps, but it doesn’t eliminate context rot on its own. A poorly ranked retrieval step can still hand a model a long, noisy context with the right fact buried in the middle of it, reproducing the exact positional bias problem context rot describes.

How do you test for context rot in a multi-tenant AI agent?

Build a small set of realistic conversations per capability, plant one checkable fact early in each, and extend them to a few increasing lengths, checking at each length whether the agent still uses that fact correctly. Run the same tests under separate tenant configurations if your product customizes behavior per customer, since a rolling summarization or truncation rule tuned for one tenant’s typical conversation length can behave differently for another tenant whose conversations run much longer.

Related reading: context engineering vs prompt engineering covers the discipline of deciding what goes into the window in the first place, upstream of the degradation this piece describes, AI agent testing covers building the eval sets this piece extends to longer conversations, AI agent observability is what to trace to catch context rot in a live conversation, and AI agent memory covers the separate, cross-session problem of what to persist once a conversation ends. See deflection rate, explained for how an unnoticed reliability problem like this one shows up in the ROI numbers you report to your own customers.

Back to Blog