This is Part 1 of a three-part series, Deep Dive: How Claude Code Works Under the Hood. In this post, we explore the agent loop. Part 2 covers the tool system and dispatch maps, and Part 3 dives into planning and subagents.
I’ve been using Claude Code daily at ZenoLab for months now — building apps, debugging gnarly Swift issues, refactoring entire codebases. And for the longest time, I treated it like a black box. I typed something, it did something smart, and I moved on.
Then I started reading through the open-source learn-claude-code project, which breaks down how agent-based coding tools actually work. What I found surprised me: the core architecture that powers something as capable as Claude Code is shockingly simple. Not simple as in “toy.” Simple as in “elegant.” The kind of simplicity that only comes from deeply understanding a problem.
Let me walk you through it.
What Is an Agent Loop?
If you’ve used Claude Code, you’ve experienced the agent loop without knowing it. You type a prompt. Claude reads your files. It edits some code. It runs your tests. It sees the test output, fixes a bug, runs the tests again — all without you doing anything after that first message.
That autonomous behavior isn’t magic. It’s a loop.
An agent loop is the repeated cycle that allows an LLM to go beyond single-shot question-and-answer and instead take actions in the real world, observe the results, and decide what to do next. It’s what turns a chatbot into an agent.
Here’s the core cycle in plain English:
- User sends a prompt (or the system provides context)
- The LLM generates a response, which may include tool calls (read a file, edit code, run a command)
- The system executes those tool calls and collects the results
- The results are written back into the conversation as new messages
- Go to step 2 — the LLM sees the updated conversation and decides what to do next
- The loop ends when the model responds without requesting any tool calls
That’s it. That is the entire architecture. Everything else — the file editing, the bash execution, the planning, the memory — is built on top of this loop.
The Write-Back Step Is the Key Insight
When I first saw this pattern, step 4 was what stopped me in my tracks. The write-back is what makes the whole thing work, and it’s the piece most people miss when they think about how AI agents function.
Here’s why it matters: after Claude Code runs a tool (say, it reads a file), the contents of that file don’t just disappear. They get appended to the conversation as a new message, as if someone had typed them in. When the LLM sees the conversation again in the next iteration, it has all the context from every previous tool call.
Think of it like this. Imagine you’re texting with a brilliant friend who can’t see your screen. You say “can you check what’s in my package.json?” They can’t do that — they’d have to ask you to paste it. But now imagine you have a helper that automatically reads the file and pastes the contents into the chat. Your friend sees it and responds intelligently. That’s the write-back.
In pseudo-code, the entire agent loop looks something like this:
messages = [{ "role": "user", "content": user_prompt }]
while True:
response = llm.call(messages=messages)
# If the model didn't ask for any tools, we're done
if not response.tool_calls:
print(response.text)
break
# Execute each tool call and write results back
for tool_call in response.tool_calls:
result = execute_tool(tool_call.name, tool_call.arguments)
messages.append({
"role": "tool",
"content": result,
"tool_call_id": tool_call.id
})
That’s roughly 15 lines of meaningful code. The core of an agent that can autonomously navigate your codebase, make edits, run tests, and iterate on failures — in 15 lines.
The Brilliance of Simplicity
When I showed this to a friend who’s a senior engineer, his first reaction was “there has to be more to it.” And yes, production Claude Code has error handling, streaming, rate limiting, authentication, and dozens of other concerns. But the fundamental loop? It never changes.
This is an important architectural insight: the core loop is stable even as the system grows more powerful. Adding a new tool — say, a tool that searches the web or creates a pull request — doesn’t require changing the loop. You just register the tool so the LLM knows about it, implement the execution logic, and the loop picks it up automatically.
I’ve seen this principle play out in my own work at ZenoLab. The best architectures I’ve built are the ones where the core is dead simple and all the complexity lives at the edges. The agent loop is a perfect example. The loop itself is trivial. The tools, the prompts, the context management — that’s where the nuance lives. But none of it requires modifying the fundamental cycle.
An Analogy That Clicked for Me
Here’s how I think about it as an indie developer: imagine you hired a brilliant assistant. Previously, you could only communicate via letters — you’d describe a problem, they’d write back with a suggestion, and you’d go implement it yourself. That’s traditional LLM usage. One prompt, one response.
Now imagine that same assistant gets a set of tools: they can open files on your computer, run terminal commands, edit code directly, and check the results. And crucially, after each action, they can see what happened and decide what to do next.
That’s the agent loop. The assistant didn’t get smarter — they got the ability to act and observe in a cycle. The intelligence was always there. The loop is what unlocks it.
This is why Claude Code feels so different from a chatbot. When I ask it to “fix the failing tests in my project,” it doesn’t just suggest what might be wrong. It reads the test file, runs the tests, sees the actual error output, traces it to the source, makes a fix, runs the tests again, and only stops when they pass. Each of those steps is one iteration of the loop.
What the Loop Looks Like in Practice
Let me trace through a real scenario. Say I’m working on one of my ZenoLab apps and I type:
“The unit tests for the SleepSessionManager are failing. Fix them.”
Here’s what happens inside the agent loop:
Iteration 1: Claude receives my prompt. It decides it needs to see the test file first. It calls the read_file tool on SleepSessionManagerTests.swift. The file contents are written back into the conversation.
Iteration 2: Claude now sees the test code. It decides to run the tests to see the actual failure. It calls bash with swift test --filter SleepSessionManager. The test output (including the failure message) is written back.
Iteration 3: Claude sees the error: “Expected 8 hours, got nil.” It reads the source file SleepSessionManager.swift to understand why the calculation returns nil. The source code is written back.
Iteration 4: Claude spots the bug — an optional that isn’t being unwrapped properly. It calls edit_file to fix the nil-coalescing logic. The edit confirmation is written back.
Iteration 5: Claude runs the tests again. They pass. The output is written back.
Iteration 6: Claude sees all tests passing. It responds with a summary of what it found and fixed. No tool calls. The loop ends.
Six iterations. Each one is the exact same cycle: the model sees the conversation, decides on an action, the system executes it, the result is written back. The loop didn’t need to know anything about Swift, about tests, about nil-coalescing. It just faithfully executed the cycle while the LLM did the thinking.
Production Considerations
If you’re thinking “okay, but real Claude Code must be more complex” — you’re right about the implementation details, but wrong about the architecture. Let me address a few things that don’t change the fundamental pattern:
Streaming: In production, you want to stream the LLM’s response token by token so the user sees output immediately. This is a presentation concern. The loop still waits for the full response (including any tool calls) before executing tools and writing back.
Error handling: What happens if a tool call fails? The error message gets written back as the tool result, and the model adapts. This is one of the most elegant aspects of the design — errors are just information. The LLM reads the error and tries something different. You don’t need special error-handling logic in the loop.
Token limits: As the conversation grows with all those write-back messages, you eventually approach the model’s context window. Production systems handle this with truncation, summarization, or context management. But the loop pattern remains identical.
Multiple tool calls: The model can request several tool calls in a single response (read three files at once, for example). The system executes all of them and writes back all the results. It’s a minor optimization, not a structural change.
Why This Matters for Your Own Projects
Understanding the agent loop has changed how I think about building AI-powered features in my own apps. The pattern is universal. If you’re building any system where an LLM needs to take actions and respond to results, this is the architecture.
At ZenoLab, I’ve started using a simplified version of this loop for in-app features — things like an AI assistant that can query a user’s data, generate reports, and iterate on formatting. The core is always the same: prompt, respond, execute, write back, repeat.
The agent loop is proof that the most powerful patterns in software are often the simplest. A while loop, a function call, and a list of messages. That’s the heart of Claude Code.
What’s Next
In the next post, we’ll look at what happens inside execute_tool — the tool system and dispatch maps that give Claude Code its hands. How does it safely read files, edit code, and run bash commands without destroying your system? That’s where things get interesting.
Continue to Part 2: Tool Use & Dispatch Maps — How Claude Code Executes Your Commands