All posts

Deep Dive: How Claude Code Works Under the Hood (Part 7) — Error Recovery That Actually Works

claude-codedeep-diveerror-handlingAI-agentreliability

This is Part 7 of the series “Deep Dive: How Claude Code Works Under the Hood.” We’ve been dissecting the architecture that powers Claude Code — from the core agent loop to permission systems, context management, and tool execution. Today we tackle something every production system must get right: what happens when things go wrong.

Failures Are Signals, Not Catastrophes

Here’s something I had to unlearn when building agent-driven workflows at ZenoLab: most failures aren’t true breakdowns. They’re signals pointing toward alternative paths.

When Claude Code hits a rate limit, it doesn’t crash. When output gets truncated mid-sentence, it doesn’t lose its train of thought. When the context window overflows, it doesn’t start hallucinating. Each of these failures is expected, classified, and routed to a specific recovery strategy.

This mindset shift — from “error = broken” to “error = redirect” — is what separates toy agents from production-grade systems. And it’s baked into every layer of Claude Code’s architecture.

The Three Recoverable Failure Types

Claude Code’s error recovery isn’t a single catch-all mechanism. It classifies failures into three distinct categories, each with its own recovery path. Let me walk through them.

1. Truncation: Output Limit Hit

Large language models have a maximum output length per response. When Claude Code is generating a long file edit or a detailed explanation, it can hit this ceiling mid-stream. The output just… stops.

The naive approach would be to throw an error and ask the user to retry. Claude Code does something smarter: it detects the truncation and injects a continuation reminder into the next turn of the conversation.

[System]: Your previous response was truncated at the output limit.
Please continue from where you left off. The last content was:
"...function processQueue(items) {"

The model picks up exactly where it stopped. From the user’s perspective, the output appears seamless — maybe with a brief pause. No data lost, no context confused.

This works because truncation is a predictable failure. The agent knows the output limit exists, can detect when it’s been hit (the response ends without a proper stop signal), and has a well-defined recovery: prompt the model to continue.

2. Context Overflow: Prompt Exceeds Limits

This is the big one. As conversations grow — tool outputs accumulate, file contents get read, error logs pile up — the total context can exceed the model’s input limit. Unlike truncation (which is about output), context overflow means the input to the model is too large.

Claude Code handles this through auto-compaction. When the prompt size approaches the limit, the system triggers a compression pass:

if (contextTokens > MAX_CONTEXT * COMPACTION_THRESHOLD) {
  conversation = await compactContext(conversation);
}

Compaction summarizes older turns while preserving recent context and critical state (like current file paths, task objectives, and tool results). The model essentially gets a “previously on…” summary instead of the full transcript.

I covered compaction in detail in an earlier post in this series, but the key insight here is that it’s an error recovery mechanism. It fires automatically when a specific failure condition (context overflow) is detected. The user never has to think about it.

3. Transport Errors: Timeouts and Rate Limits

Network requests fail. APIs return 429 (rate limited) or 503 (service unavailable). Connections time out. These are the bread-and-butter failures of any networked application, and Claude Code handles them with exponential backoff with jitter.

async function retryWithBackoff(operation, maxRetries = 5) {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      return await operation();
    } catch (error) {
      if (!isTransientError(error)) throw error;

      const baseDelay = Math.min(1000 * Math.pow(2, attempt), 30000);
      const jitter = Math.random() * baseDelay * 0.5;
      await sleep(baseDelay + jitter);
    }
  }
  throw new MaxRetriesExceededError();
}

The exponential part means each retry waits longer: 1s, 2s, 4s, 8s, up to a cap. The jitter adds randomness so that when multiple requests fail simultaneously (say, during an API rate limit), they don’t all retry at exactly the same moment and cause another spike.

Notice the isTransientError check — not all errors get retried. A 401 (unauthorized) is not transient. A 404 (not found) won’t magically resolve itself. Only errors that are genuinely temporary trigger the retry loop.

The Classify-Then-Act Pattern

The unifying design principle across all three recovery types is what I call classify-then-act. Every error passes through a classification stage before any recovery logic runs.

function classifyError(error) {
  if (error.type === 'truncation') return 'TRUNCATION';
  if (error.type === 'context_overflow') return 'CONTEXT_OVERFLOW';
  if (isTransientError(error)) return 'TRANSPORT';
  if (error.type === 'tool_failure') return 'TOOL_FAILURE';
  return 'UNRECOVERABLE';
}

function recover(error, state) {
  const category = classifyError(error);

  switch (category) {
    case 'TRUNCATION':
      return injectContinuationReminder(state);
    case 'CONTEXT_OVERFLOW':
      return compactAndRetry(state);
    case 'TRANSPORT':
      return retryWithBackoff(state);
    case 'TOOL_FAILURE':
      return reportAndContinue(state);
    case 'UNRECOVERABLE':
      return escalateToUser(state);
  }
}

This is fundamentally different from wrapping everything in a generic try/catch. The classification makes recovery visible — you can log it, monitor it, and reason about it. There are no hidden catch blocks silently swallowing errors. Every recovery action is an explicit state transition in the agent loop.

When I adopted this pattern in my own projects at ZenoLab, debugging agent failures went from “something went wrong somewhere” to “we hit a TRANSPORT error on attempt 3 of 5, backing off for 4.2 seconds.” Night and day difference.

Recovery Counters: Preventing Infinite Loops

Here’s a trap that catches a lot of agent builders: what happens when recovery itself fails? If truncation recovery triggers another truncation, and that triggers another, you get an infinite loop burning tokens and time.

Claude Code prevents this with recovery counters — simple integers that track how many times each recovery type has fired within a given session or task.

const recoveryCounters = {
  truncation: 0,
  contextOverflow: 0,
  transport: 0,
};

function attemptRecovery(category, state) {
  recoveryCounters[category]++;

  if (recoveryCounters[category] > MAX_RECOVERIES[category]) {
    return escalateToUser(state, category);
  }

  return recoveryStrategies[category](state);
}

Each category has its own maximum. Transport errors might get 5 retries (they’re cheap — just waiting and retrying a network call). Context overflow compaction might get 2 attempts (each one is expensive and lossy). Truncation continuations might get 3 tries before the system says “this output is too large for me to complete in pieces.”

Budget-Based Retry Limits

Recovery counters are per-category for a reason. Treating all errors the same way is one of the most common mistakes in agent error handling.

Consider the cost profile:

Error TypeCost Per RetryTypical Budget
TransportLow (just network wait)5 retries
TruncationMedium (model inference)3 retries
Context OverflowHigh (compaction + inference)2 retries

A transport retry costs a few seconds of wall time. A truncation retry costs a full model inference call (tokens, latency, money). A context overflow retry costs a compaction pass plus a model inference call — it’s the most expensive recovery operation.

Setting the same retry limit for all three would either over-retry expensive operations (wasting resources) or under-retry cheap ones (giving up too early on easily recoverable failures).

This is budget-based thinking applied to error handling. Every retry spends resources — time, tokens, API calls, user patience. The budget should match the cost.

Why Identical Retry Logic Is Wrong

I want to hammer this point because I’ve seen it in too many agent implementations (including my own early attempts at ZenoLab). The pattern looks like this:

// DON'T DO THIS
async function handleError(error, retries = 3) {
  if (retries <= 0) throw error;
  await sleep(1000);
  return retry(retries - 1);
}

Three retries, one-second delay, same logic regardless of error type. This fails in multiple ways:

  • Transport errors need exponential backoff, not fixed delays. Hammering a rate-limited API every second makes things worse.
  • Truncation doesn’t need a delay at all — it needs a continuation prompt immediately.
  • Context overflow needs compaction first, and if compaction already ran, retrying without changing anything is pointless.

The classify-then-act pattern ensures each error type gets the recovery strategy it actually needs.

Visible State, Not Hidden Catch Blocks

One design choice that stands out in Claude Code’s error handling is the emphasis on visibility. Recovery actions are not buried in nested try/catch blocks three layers deep. They’re represented as explicit states in the agent loop.

When the agent is recovering from truncation, that’s a visible state: CONTINUING_AFTER_TRUNCATION. When it’s waiting on a transport retry, that’s RETRYING_TRANSPORT_ATTEMPT_3. When compaction fires, it’s COMPACTING_CONTEXT.

This matters for three reasons:

  1. Debugging: When something goes wrong with the recovery itself, you can see exactly which state the agent was in.
  2. Monitoring: You can track recovery rates across sessions and identify systemic issues (e.g., “truncation recovery fires 40% of the time — maybe our prompts are too verbose”).
  3. User communication: The agent can tell the user what’s happening. “I’m continuing my previous response” is better than a mysterious pause.

Completing the Hardening Stage

This post on error recovery completes what I think of as the “hardening” stage of agent design. Over the last few posts, we’ve covered:

  • Permission systems that prevent dangerous operations
  • Context management that keeps the agent working within model limits
  • Tool execution with sandboxing and validation
  • Error recovery that classifies, budgets, and visibly handles failures

These four pillars turn a basic prompt-tool-response loop into something you can actually deploy and trust. The agent won’t silently corrupt your codebase, won’t spin forever on a failed API call, won’t lose context mid-task, and won’t crash on predictable failure modes.

From here, we move into higher-level features: task management, background execution, and the systems that let Claude Code handle complex, multi-step work. The foundation is solid — now we build on it.

What’s Next

In Part 8, we’ll explore how Claude Code manages persistent tasks — not just flat to-do lists, but structured task DAGs (directed acyclic graphs) with dependencies, status tracking, background execution, and even cron scheduling. This is where the agent starts to feel less like a chatbot and more like a project manager that never sleeps.

More from the studio.

Back to blog