LangGraph Reliability

LangGraph observability for graphs that run in production

LangGraph gives you the three things a real agent needs — cycles, durable state, and human-in-the-loop interrupts — and each one is a failure mode a linear chain never had. Graphs loop until the recursion limit. State corrupts in one node and surfaces three nodes later. A thread parked on an interrupt looks identical to a thread that finished. Reliability starts with tracing the run as a graph, then turning those traces into numbers your team and your customers can act on.

The problem

Cycles and state are the point, and the risk

LangGraph exists because useful agents are not straight lines. They loop, they branch on what a tool returned, they pause for a human and resume hours later. That expressiveness is exactly why a graph run is harder to reason about after the fact than a chain: the execution path is decided at runtime, and the code alone will not tell you which path a given run took.

Four failure modes show up again and again once a graph is live:

Runaway cycles. A conditional edge routes back into a tool-calling node that keeps asking for the same tool. The run ends at recursion_limit with a GraphRecursionError, having spent real money. The limit caught it; nothing explains it.

State corruption between nodes. Reducers append and merge rather than overwrite, so a node returning a slightly wrong shape does not fail — it quietly poisons state for every node after it.

Routing that is wrong but not broken. A conditional edge sends the run down a valid branch that was the wrong branch for this input. The graph completes successfully. The answer is wrong.

Non-idempotent partial runs. A node with a real side effect — a refund, an email, a write — succeeds, a later node fails, and the thread resumes from its checkpoint and does it again. Durable execution replays the graph, not the consequences.

What to capture

Six things worth tracing on a graph run

Capture these and a failed thread stops being something you reproduce by hand from a checkpoint.

Node-level spans

One span per node execution, with latency and status, parented to the graph run. The unit that matters is the node, not the LLM call — a slow graph is usually one slow node, not a uniformly slow model.

The path actually taken

Which conditional edge fired at each branch and what the routing function saw. Two runs of the same graph on similar input can take different paths, and the difference between them is the bug report.

State deltas per node

What each node returned, not just the final state. With appending reducers, the node that corrupted state and the node that visibly failed are rarely the same node.

Super-steps against the recursion limit

Steps consumed, and which nodes repeated. A run that ends at the ceiling is a truncated run; a run that ends at 80% of the ceiling is a warning you still have time to act on.

Interrupts and resumes

When a thread pauses for human input, how long it waits, and whether it ever resumes. Threads abandoned mid-interrupt are invisible in most dashboards and are pure lost work.

Tokens and cost per thread

Cost attributed to a thread and a customer, not just a service. Cycles make graph cost far more variable than chain cost, so per-run cost is the number that catches a regression early.

Getting data in

Start with the exporter you already have

LangGraph apps instrumented for OpenTelemetry can ship traces without graph changes. The ingest gateway reads the OTel GenAI semantic conventions and scopes every span to an org and a customer on arrival.

bash
# Any OpenTelemetry-instrumented LangGraph app can export straight to the
# ingest gateway. No change to your StateGraph, nodes, or edges.

export OTEL_EXPORTER_OTLP_ENDPOINT="https://ingest.aiagre.com"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer aig_your_key,\
x-aiagre-customer-id=cus_acme,\
x-aiagre-agent-id=support-graph"
export OTEL_SERVICE_NAME="support-graph"

python -m app.serve

Spans arrive as structured runs — prompts, tool calls, tokens, model, latency, and status — keyed by trace ID rather than flattened into logs.

Then tie each thread to an outcome

Traces explain a single run. Outcomes make runs comparable, and they are the input to deflection rate, resolution rate, and cost per resolution.

typescript
// One thread = one run. Tie the graph's traces to the outcome it produced.
import { init, withRun, traceChain, trackOutcome } from "@aiagre/node";

init({
  apiKey: process.env.AIAGRE_API_KEY!,
  customerId: "cus_acme",    // your end customer
  agentId: "support-graph",  // the graph this run belongs to
});

// sessionId is the LangGraph thread_id, so traces, checkpoints, and
// outcomes all line up on the same identifier.
await withRun({ sessionId: threadId }, async () => {
  const state = await traceChain({ name: "graph.invoke" }, () =>
    graph.invoke(input, { configurable: { thread_id: threadId } })
  );

  await trackOutcome({
    name: "conversation.completed",
    outcome: state.needsHuman ? "handoff_to_human" : "resolved",
    metadata: { steps: state.stepCount, interrupted: state.wasInterrupted },
  });
});

The Node SDK covers LangGraph.js and other TypeScript services; Python graphs use the OTLP path above. See AI agent reliability for how traces and outcomes fit together.

The part tools skip

Your customers will never open a graph trace

Node spans and state deltas serve your engineers. If your graph runs inside a product other companies pay for, those customers arrive with a different question: is this agent saving us money? A trace viewer does not answer that, and it was never built to.

Answering it means deriving a few numbers from the same traces — deflection rate, cost per resolution, resolution rate — and being explicit about the definitions, because none of them are standardized. Publish the formula next to the number and it survives scrutiny.

The harder half is architectural. Showing customer A their slice of your graph data, with no path by which A sees customer B's threads, means every span carries an org identity and a customer identity from ingestion, and every read is scoped the same way. Filtering a single-tenant dashboard after the fact is how tenant data leaks.

AiAgRe scopes every ingested event to an org and a customer, then issues short-lived embed tokens so each customer's dashboard reads only their own threads — inside your product, under your branding.

The full loop

Observe, prove, remediate

Tracing a graph is the first third. The other two thirds are what keep it reliable after the first incident.

Observe

Node spans, routing decisions, state deltas, interrupts, tokens, and status, tied to a thread you can replay instead of reconstruct.

Prove

Deflection rate, cost per resolution, and resolution rate computed from those traces, scoped per customer and defensible when someone asks how the number was derived.

Remediate

Cluster recurring graph failures — runaway cycles, bad branches, stalled interrupts — tighten the guardrails that matter, then confirm the numbers moved.

Checkpoints are not metrics:A checkpointer tells you what one thread's state was. It cannot tell you which conditional edge fires most often, or what a run costs per customer.

FAQs

LangGraph observability: frequently asked questions

Common questions from teams running LangGraph agents in production.

What is LangGraph observability?

LangGraph observability means recording a graph run as a graph: which nodes executed, in what order, how the state changed at each step, which conditional edge was taken and why, how many super-steps the run consumed against its recursion limit, and where it stopped. A flat list of LLM calls loses the routing decisions, and routing is where most graph bugs live.

Why does my LangGraph agent hit GraphRecursionError?

Because a conditional edge keeps routing back into a node that never satisfies its exit condition — usually a tool-calling loop where the model asks for the same tool again after an unhelpful result. LangGraph stops the run at recursion_limit, which protects your bill but tells you nothing about the cycle. Raising the limit is the wrong fix roughly every time; the trace of which nodes repeated is what points at the real one.

How do I debug LangGraph state between nodes?

Record the state delta each node returns, not just the final state. Reducers like add_messages append rather than overwrite, so a node that returns a slightly wrong shape corrupts state quietly and the symptom appears several nodes later. Per-node deltas turn that into a specific node and a specific key instead of a whole-graph mystery.

Are LangGraph checkpoints enough for production monitoring?

Checkpointers are excellent for durability and time travel: they let a thread resume after a crash or an interrupt, and let you replay from a past state. They are not aggregate observability. A checkpoint answers what this one thread's state was; it does not tell you the p95 latency of your routing node, which conditional edge fires most often, or what a graph run costs per customer.

How do I monitor a LangGraph app in production?

If your app already emits OpenTelemetry, point the OTLP exporter at an ingest endpoint that understands the GenAI semantic conventions — you get node-level traces without touching graph code. Then record a business outcome per thread: resolved, handed off to a human, abandoned. Traces explain a run; outcomes make runs comparable.

Can my customers see their own LangGraph metrics?

Only if the data was scoped per customer at write time. AiAgRe tags every ingested span with an org identity and a customer identity, then issues short-lived embed tokens so each customer's dashboard reads only their own graph runs, cost, and resolution numbers — rendered inside your product rather than on a vendor portal.

Ship a graph you can prove works

Request access and we'll help you instrument your first graph and scope a customer-facing dashboard.