All posts

Deep Dive: Planning & Subagents — How Claude Code Stays Focused on Complex Tasks

claude-codeaiplanningsubagentsTodoWritecontext-managementdeep-dive

This is Part 3 of a three-part series, Deep Dive: How Claude Code Works Under the Hood. Part 1 covered the agent loop and Part 2 explored the tool system and dispatch maps.

In Parts 1 and 2, we saw how simple the core of an AI agent really is: a loop that calls tools and writes results back, with a dispatch map routing tool names to handlers. If you stopped there, you’d have a working agent. But you’d also have an agent that falls apart on anything complex.

Here’s the problem: ask Claude Code to “refactor the authentication module, update all the tests, and fix the broken CI pipeline,” and the naive loop has to hold all of that in its head at once. As it works through the first task, the conversation fills up with file contents, test outputs, and error messages. By the time it gets to the CI pipeline, the original instructions are buried under thousands of tokens of context. The model starts to drift. It forgets what it was doing. It loses the thread.

This is not a theoretical concern — it’s the central challenge of building useful AI agents. And Claude Code has two elegant solutions: structured planning and subagents. Let me walk you through both.

The Drift Problem

If you’ve used any LLM for extended conversations, you’ve experienced drift firsthand. You start talking about database optimization, and 20 messages later the model is happily discussing CSS animations because the conversation wandered there naturally.

For a chatbot, drift is annoying. For an autonomous agent that’s editing your codebase, drift is dangerous. Imagine Claude Code starts refactoring your auth module, gets distracted by an interesting type error in an unrelated file, “fixes” it by changing an API contract, and now half your application is broken. Not because the model is dumb — because it lost focus.

The root cause is the context window. Every LLM has a finite amount of text it can consider at once. As the agent loop runs and tool results pile up, the early messages (including the user’s original instructions) get pushed further back. The model pays decreasing attention to them. It starts optimizing for whatever is most recent and salient in the conversation, not what the user actually asked for.

This is why a naive agent loop works great for simple tasks (“read this file and explain it”) but struggles with complex, multi-step work. You need mechanisms to keep the agent on track.

TodoWrite: A Structured Plan the Agent Can Follow

Claude Code’s first defense against drift is a tool called TodoWrite. It’s deceptively simple: a tool that lets the agent create and manage a structured task list during execution.

When Claude Code receives a complex request, one of the first things it does is break the work into discrete tasks and write them down using TodoWrite:

{
  "tool": "TodoWrite",
  "arguments": {
    "todos": [
      { "id": "1", "task": "Read and understand current auth module", "status": "in_progress" },
      { "id": "2", "task": "Refactor auth module to use new token pattern", "status": "pending" },
      { "id": "3", "task": "Update unit tests for refactored auth", "status": "pending" },
      { "id": "4", "task": "Fix CI pipeline configuration", "status": "pending" }
    ]
  }
}

This isn’t just for show. The todo list persists in the conversation and serves as a structural anchor. When the model sees this list in its context, it’s reminded of the full scope of work and where it currently stands. It’s the difference between working from memory and working from a checklist.

The Single-Task Constraint

Here’s the clever part: the system enforces that only one task can have in_progress status at a time. The agent must complete or explicitly set aside the current task before starting the next one.

Why does this matter? Because LLMs are easily tempted by tangents. Without this constraint, the model might start three tasks simultaneously, interleave their tool calls, and create a confused mess of half-completed changes. The single-task constraint forces sequential focus.

When the agent finishes a task, it updates the todo list:

{
  "tool": "TodoWrite",
  "arguments": {
    "todos": [
      { "id": "1", "task": "Read and understand current auth module", "status": "completed" },
      { "id": "2", "task": "Refactor auth module to use new token pattern", "status": "in_progress" },
      { "id": "3", "task": "Update unit tests for refactored auth", "status": "pending" },
      { "id": "4", "task": "Fix CI pipeline configuration", "status": "pending" }
    ]
  }
}

This creates a clear narrative in the conversation: here’s what’s done, here’s what I’m doing now, here’s what’s left. The model can always orient itself by looking at the most recent todo state.

I’ve started using a similar pattern in my own planning when I work on ZenoLab projects. Before I start a complex feature, I write out the task list in my project management tool. Not because I’ll forget — but because the act of structuring the work prevents me from getting pulled into rabbit holes. The same psychology applies to LLMs, except for them it’s not psychology, it’s attention mechanics.

Nag Injection: The Automated Course Correction

TodoWrite helps the agent plan, but what happens when it starts drifting despite having a plan? This is where nag injection comes in.

Nag injection is a system-level mechanism that automatically inserts reminder messages into the conversation when the agent appears to be going off track. The concept is straightforward: if the agent has been running for a while without updating its todo list, or if it’s making tool calls that don’t seem related to the current in_progress task, the system injects a message like:

“Reminder: Your current task is ‘Refactor auth module to use new token pattern.’ You have 3 remaining tasks. Please stay focused on the current task or update your plan if priorities have changed.”

The model sees this as a system message in the next iteration of the agent loop. It’s not interrupting the model — it’s adding information to the context that nudges the model back on course.

This is one of those ideas that sounds almost too simple to work. But it works because of how LLMs process context: recent messages have strong influence on the next generation. A well-timed reminder in the right position can completely redirect the model’s behavior.

Think of it like a gentle tap on the shoulder. You’re deep in a debugging session, you’ve been chasing a tangent for 20 minutes, and a colleague says “hey, weren’t you working on the auth refactor?” Instantly, you reorient. Nag injection is that colleague, automated.

Subagents: Disposable Scratchpads for Complex Subtasks

Planning and nag injection solve the focus problem. But there’s another problem that’s just as critical: context pollution.

Every time the agent reads a file, runs a command, or gets test output, those results accumulate in the conversation. After 50 tool calls, the conversation might contain thousands of lines of code, dozens of error messages, and pages of terminal output. Most of it is no longer relevant — it was useful for the specific step where it was generated, but now it’s just noise consuming precious context window space.

This is where subagents come in, and honestly, this is my favorite architectural pattern in the entire system.

A subagent is a child agent loop that runs with its own isolated context. The parent agent spawns a subagent for a specific subtask, the subagent does its work in a fresh conversation, and when it’s done, only the summary comes back to the parent. Everything else — all the intermediate file reads, test runs, error messages — is discarded.

Here’s the pattern:

Parent Agent Context:
  [user request]
  [todo list]
  [high-level progress]

    ↓ spawns subagent for "refactor auth module"

    Subagent Context (isolated):
      [specific instructions from parent]
      [reads auth module... 200 lines of code]
      [makes edits... confirmation]
      [runs tests... test output]
      [fixes failing test... more output]
      [all tests pass]

    ↑ returns summary: "Refactored auth module to use new token
       pattern. Updated 3 files. All 12 tests passing."

Parent Agent Context:
  [user request]
  [todo list]
  [summary from subagent]  ← only this comes back
  [continues to next task]

The parent never sees the 200 lines of code, the test output, or the intermediate failures. It gets a clean summary and moves on. Its context stays focused and uncluttered.

Why Context Isolation Matters

Let me put some numbers on this. Say a complex task involves 80 tool calls, each returning an average of 100 lines of content. That’s 8,000 lines of context accumulated — easily 40,000+ tokens. In a model with a 200k token context window, that’s 20% of your budget consumed by intermediate results from earlier subtasks that are no longer relevant.

With subagents, those 80 tool calls might be split across 4 subagents of 20 calls each. Each subagent works within its own clean 5,000-token context. The parent sees four summaries totaling maybe 500 tokens. That’s a 99% reduction in context pollution.

This is why Claude Code stays fast and coherent even on large, complex tasks. It’s not that the model is somehow immune to context limits — it’s that the architecture actively prevents context waste.

The Parent-Child Model

The relationship between parent and child agents follows a clean contract:

  1. Parent sends specific instructions — not “do everything,” but “refactor the auth module to use token-based authentication, here are the relevant files”
  2. Child executes autonomously — it has its own agent loop, its own tool access, and full autonomy to read files, make edits, run tests
  3. Child returns a summary — a concise description of what was done, what changed, and the final state
  4. Parent integrates the summary — updates the todo list, decides the next step
  5. Child’s context is discarded — all those intermediate tool results vanish

This is analogous to delegation in a real team. When I delegate a task to a contractor at ZenoLab, I don’t want a minute-by-minute log of everything they did. I want to know: what changed, does it work, are there any issues I should know about? That’s exactly what the subagent returns.

How It All Fits Together

Let’s trace through a complex request to see all these mechanisms working in concert.

I type: “Refactor the sleep tracking module to use the new HealthKit API, update all tests, and make sure the CI pipeline passes.”

Step 1: Planning. Claude Code creates a todo list with TodoWrite:

  • Read current sleep tracking implementation (in_progress)
  • Refactor to new HealthKit API (pending)
  • Update unit tests (pending)
  • Update integration tests (pending)
  • Verify CI pipeline (pending)

Step 2: Research subtask. The parent spawns a subagent to read and analyze the current implementation. The subagent reads 8 files, traces the data flow, and returns a summary: “Current implementation uses HKCategoryType for sleep analysis. Three main classes involved: SleepSessionManager, SleepDataStore, HealthKitBridge. 47 unit tests, 12 integration tests.”

Step 3: Refactoring subtask. The parent spawns another subagent with specific instructions based on the research summary. This subagent reads the files, makes edits, runs tests, iterates on failures — all in its own context. It returns: “Refactored all three classes to use new HKSleepAnalysis API. Updated 312 lines across 5 files. Unit tests: 43 passing, 4 need updates (expected, due to API changes).”

Step 4: Test update subtask. Another subagent fixes the 4 failing tests and updates integration tests. Returns: “All 47 unit tests passing. Updated 12 integration tests, all passing.”

Step 5: CI verification. Final subagent checks the CI configuration and runs a local verification. Returns: “CI pipeline configuration is compatible. Local build succeeds.”

Step 6: Completion. The parent marks all tasks complete and responds to me with a clean summary.

Throughout this entire process, the parent’s context stayed lean. It held the todo list, the user request, and five concise summaries. Meanwhile, the subagents did the heavy lifting in isolated contexts that were discarded after use.

And nag injection? If at any point a subagent started drifting (say, the refactoring subagent noticed an unrelated performance issue and started optimizing it), the system would inject a reminder to stay focused on the HealthKit refactoring. The subagent would course-correct and continue.

Practical Implications

Understanding these mechanisms has changed how I interact with Claude Code:

  1. I give complex instructions confidently. Knowing that Claude Code has planning and context isolation, I’m not afraid to ask for multi-step tasks. It won’t lose the thread.

  2. I write better prompts. When I know the agent will create a todo list, I structure my requests to map naturally to discrete tasks. “Do A, then B, then C” works better than a vague paragraph.

  3. I trust the process. When Claude Code is working on step 3 of a 5-step task, I don’t worry that it’s forgotten steps 4 and 5. The todo list has them. The nag injection will enforce them.

  4. I understand the limits. If a task is so complex that even with subagents the context gets overwhelmed, I know to break it into separate conversations rather than fighting the architecture.

These insights apply beyond Claude Code. If you’re building any multi-step AI workflow — whether it’s an automated code review pipeline, a data analysis agent, or an AI assistant in your app — the principles of structured planning and context isolation will make your system dramatically more reliable.

The Full Picture

Across this three-part series, we’ve seen how Claude Code works from the inside out:

  • The agent loop (Part 1) provides the fundamental cycle: prompt, respond, execute tools, write back results, repeat until done
  • The tool system (Part 2) gives the agent safe, structured ways to interact with the world through dispatch maps and purpose-built tools with security layers
  • Planning and subagents (Part 3) keep the agent focused and context-efficient on complex, multi-step tasks

What strikes me most is how each layer is independently simple. The loop is a while loop. The dispatch map is a dictionary. TodoWrite is a list with status fields. Subagents are just the same loop running in a child process. The power comes from composing these simple pieces together.

That’s a lesson I keep relearning as an indie developer at ZenoLab: the best systems aren’t the ones with the most clever code. They’re the ones with the simplest pieces, combined thoughtfully.

If you’re building AI-powered tools or thinking about agent architectures, I hope this series gave you a concrete mental model. The principles here — the loop, the dispatch map, structured planning, context isolation — aren’t specific to Claude Code. They’re patterns you can apply to any system where an LLM needs to do real work in the real world.

Happy building.


This is the final post in the Deep Dive series. If you’re new to Claude Code and want to get started, check out Getting Started with Claude Code. If you want to see how I use it in practice for building apps, read Vibe Coding with Claude Code.

More from the studio.

Back to blog