All posts

Deep Dive: How Claude Code Works Under the Hood (Part 9) — Agent Teams & Coordination Protocols

claude-codedeep-divemulti-agentagent-teamsautonomous-agents

This is Part 9 of the series “Deep Dive: How Claude Code Works Under the Hood.” Last time we explored task systems and background execution — how Claude Code manages persistent work with DAGs, daemon threads, and cron scheduling. Now we get to something I find genuinely exciting: multiple agents working together as a coordinated team.

From Disposable Subagents to Persistent Teammates

If you’ve used Claude Code’s subagent feature (which we covered earlier in this series), you’ve seen the basic pattern: the main agent spawns a child agent, gives it a focused task, collects the result, and the child disappears. It’s like calling a function — useful, but stateless.

Agent teams are a fundamentally different concept. Instead of disposable workers, you get persistent teammates — independent agent loops that maintain their own context, have their own inboxes, and coordinate through structured protocols. They don’t disappear after one task. They stick around, pick up new work, and communicate.

This is the difference between hiring a contractor for a single afternoon and having a team that works together over weeks. The coordination overhead is higher, but the capabilities are in a different league.

The .team/ Directory

Everything starts with a directory. Claude Code’s team system is anchored in a .team/ directory at the project root:

.team/
├── config.json            # Team configuration
├── agents/
│   ├── frontend/
│   │   ├── identity.json  # Agent identity and capabilities
│   │   └── inbox.jsonl    # Message queue
│   ├── backend/
│   │   ├── identity.json
│   │   └── inbox.jsonl
│   └── reviewer/
│       ├── identity.json
│       └── inbox.jsonl
└── shared/
    └── taskboard.json     # Shared task board

The config.json defines team composition — how many agents, what roles, what capabilities each one has. Each agent gets its own subdirectory with an identity file and a JSONL inbox for receiving messages.

This file-based architecture is a deliberate choice. Files are inspectable, debuggable, and survive process restarts. You can cat an agent’s inbox to see what messages it received. You can edit config.json to add or remove team members. There’s no opaque message broker or in-memory queue that evaporates when the process dies.

TeammateManager: Spawning Independent Loops

The TeammateManager is the orchestrator. It reads the team config, spawns each agent as an independent loop, and manages their lifecycles.

class TeammateManager {
  constructor(config) {
    this.agents = new Map();
    this.config = config;
  }

  async spawnTeam() {
    for (const agentConfig of this.config.agents) {
      const agent = new AgentLoop({
        identity: agentConfig.identity,
        inbox: agentConfig.inboxPath,
        capabilities: agentConfig.capabilities,
      });

      this.agents.set(agentConfig.name, agent);
      agent.start(); // Non-blocking — runs its own loop
    }
  }

  getAgent(name) {
    return this.agents.get(name);
  }
}

Each spawned agent runs its own event loop independently. It reads from its inbox, processes messages, executes tasks, writes results. The TeammateManager doesn’t micromanage — it spawns and monitors, but the agents are autonomous within their roles.

This is architecturally similar to microservices. Each agent is an independent process (or at least an independent async loop) with its own state and communication channel. The coupling between agents is through messages, not shared memory.

File-Based Messaging

Communication between agents uses the simplest protocol imaginable: append a JSON line to the recipient’s inbox file.

function sendMessage(recipientName, message) {
  const inboxPath = `.team/agents/${recipientName}/inbox.jsonl`;
  const line = JSON.stringify({
    id: generateId(),
    from: currentAgent.name,
    timestamp: Date.now(),
    type: message.type,
    payload: message.payload,
  }) + '\n';

  fs.appendFileSync(inboxPath, line);
}

And reading messages:

function readNewMessages(agent) {
  const lines = fs.readFileSync(agent.inboxPath, 'utf-8')
    .split('\n')
    .filter(Boolean)
    .map(JSON.parse);

  const newMessages = lines.filter(m => m.id > agent.lastReadId);
  agent.lastReadId = lines[lines.length - 1]?.id ?? agent.lastReadId;

  return newMessages;
}

JSONL (JSON Lines) is perfect here. Each message is a single line, so appending is atomic on most file systems. No corruption risk from concurrent writes. No need for file locking in practice, because appends to separate lines don’t conflict.

The simplicity of this is something I deeply appreciate as an indie developer. No Redis. No RabbitMQ. No gRPC. Just files. You can debug agent communication with tail -f .team/agents/backend/inbox.jsonl. That’s it.

Teammate Lifecycle

Each agent follows a defined lifecycle:

spawn → WORKING → IDLE → WORKING → ... → SHUTDOWN
  • spawn: Agent is created and initialized with its identity and capabilities
  • WORKING: Agent is actively processing a task or message
  • IDLE: Agent has no pending work and is waiting for new messages
  • SHUTDOWN: Agent has been told to stop and is cleaning up

The transitions are event-driven:

async function agentLifecycle(agent) {
  agent.status = 'WORKING'; // Initial setup tasks

  while (agent.status !== 'SHUTDOWN') {
    const messages = readNewMessages(agent);
    const tasks = getClaimableTasks(agent);

    if (messages.length > 0 || tasks.length > 0) {
      agent.status = 'WORKING';
      await processWork(agent, messages, tasks);
    } else {
      agent.status = 'IDLE';
      await waitForWork(agent, IDLE_TIMEOUT);

      if (agent.idleDuration > IDLE_SHUTDOWN_THRESHOLD) {
        agent.status = 'SHUTDOWN';
      }
    }
  }

  await gracefulShutdown(agent);
}

The IDLE state isn’t just “doing nothing.” It’s an active wait with a timeout. If an agent sits idle for 60 seconds with no new messages and no claimable tasks, it transitions to SHUTDOWN. This prevents zombie agents from consuming resources indefinitely.

Structured Protocols: Request-Response with Tracking

Simple fire-and-forget messaging works for notifications, but complex coordination needs request-response semantics. Claude Code implements this with tracking IDs:

// Requester sends
sendMessage('backend', {
  type: 'REQUEST',
  payload: {
    trackingId: 'req_abc123',
    action: 'generate_api_endpoint',
    params: { resource: 'users', methods: ['GET', 'POST'] }
  }
});

// Responder replies
sendMessage('frontend', {
  type: 'RESPONSE',
  payload: {
    trackingId: 'req_abc123',
    status: 'completed',
    result: { endpoint: '/api/users', schema: '...' }
  }
});

The trackingId links the response back to the original request. The requesting agent can await a response with a specific tracking ID, implement timeouts, and handle cases where the response never arrives.

This is a minimal implementation of what enterprise systems do with correlation IDs. But it’s enough for agent coordination without introducing a full message broker.

Shutdown Protocol: The Handshake

Shutting down an agent isn’t as simple as killing the process. If an agent is mid-way through writing a file, killing it could leave corrupted state. Claude Code implements a shutdown handshake:

async function initiateShutdown(agentName) {
  // Step 1: Send shutdown request
  sendMessage(agentName, {
    type: 'SHUTDOWN_REQUEST',
    payload: { reason: 'idle_timeout', deadline: Date.now() + 10000 }
  });

  // Step 2: Wait for acknowledgment
  const ack = await waitForMessage({
    from: agentName,
    type: 'SHUTDOWN_ACK',
    timeout: 10000,
  });

  if (ack) {
    // Agent confirmed it's safe to stop
    removeAgent(agentName);
  } else {
    // Force shutdown after deadline
    forceRemoveAgent(agentName);
  }
}

The agent receives the shutdown request, finishes any in-progress file writes, flushes its state, sends an acknowledgment, and then stops. This prevents half-written files and inconsistent state.

It’s a small detail, but it’s the difference between a system that works in demos and one that works in production. At ZenoLab, I learned this the hard way — killing agents mid-write corrupted task state files twice before I implemented proper shutdown handshakes.

Plan Approval: Gating Risky Changes

Not everything should be autonomous. Some changes — deleting files, modifying configuration, pushing to production branches — need human review.

Claude Code’s team system includes a plan approval gate:

async function executeRiskyAction(agent, action) {
  if (requiresApproval(action)) {
    const plan = {
      agent: agent.name,
      action: action.type,
      targets: action.targets,
      reasoning: action.reasoning,
    };

    sendMessage('coordinator', {
      type: 'APPROVAL_REQUEST',
      payload: plan,
    });

    const decision = await waitForApproval(plan.id);

    if (decision.approved) {
      return executeAction(action);
    } else {
      return { status: 'blocked', reason: decision.reason };
    }
  }

  return executeAction(action);
}

The coordinator (which might be the main agent or the human user) reviews the plan and approves or rejects it. This puts a human checkpoint in the automation loop exactly where it’s needed — at the risky boundaries — without slowing down routine work.

Autonomous Agents and the Shared Task Board

The most advanced mode is fully autonomous operation: agents self-organizing around a shared task board.

The taskboard.json in the shared directory contains all pending tasks. Any agent can scan it:

function autoClaimTask(agent) {
  const board = readTaskBoard();

  const claimable = board.tasks.filter(task =>
    task.status === 'pending' &&
    task.owner === null &&
    task.dependencies.every(d => isCompleted(d, board)) &&
    matchesCapabilities(task, agent.capabilities)
  );

  if (claimable.length > 0) {
    const task = claimable[0];
    task.status = 'in_progress';
    task.owner = agent.name;
    writeTaskBoard(board);
    return task;
  }

  return null;
}

The auto-claim logic is elegant: find tasks that are pending, unowned, have all dependencies met, and match this agent’s capabilities. Claim the first one. This creates emergent work distribution — frontend tasks naturally flow to the frontend agent, backend tasks to the backend agent, without explicit assignment.

Identity Re-Injection After Context Compression

Here’s a subtle but critical detail. When an agent’s context gets compressed (remember auto-compaction from the error recovery post?), it might lose its sense of identity. The model forgets “I am the frontend agent specializing in React components.”

Claude Code solves this by re-injecting the agent’s identity after every compaction:

async function compactAgentContext(agent) {
  agent.conversation = await compactContext(agent.conversation);

  // Re-inject identity
  agent.conversation.unshift({
    role: 'system',
    content: `You are ${agent.identity.name}. ${agent.identity.description}. 
              Your capabilities: ${agent.identity.capabilities.join(', ')}.
              You are part of a team. Check your inbox for messages.`
  });
}

Without this, a compacted agent might start trying to do work outside its role, or forget to check its inbox entirely. The identity re-injection keeps agents on track across long sessions.

The 60-Second Idle Timeout

I mentioned this briefly, but it deserves its own callout. When an agent has been idle for 60 seconds — no messages, no claimable tasks — it initiates shutdown.

Why 60 seconds? It’s a balance:

  • Too short (10 seconds): Agents shut down during brief lulls and have to be re-spawned constantly. Spawn overhead adds up.
  • Too long (10 minutes): Idle agents consume memory and context capacity for no reason.
  • 60 seconds: Long enough to survive brief pauses between tasks, short enough to not waste resources.

This is configurable in config.json, but the default reflects real-world usage patterns where tasks typically arrive in clusters with brief gaps between them.

What’s Next

In the final post of this series (Part 10), we’ll cover the last two architectural pieces: worktree isolation for preventing file conflicts between concurrent agents, and MCP (Model Context Protocol) for extending Claude Code’s capabilities through plugins. Then we’ll wrap up with a retrospective on the full architecture — from the simple agent loop in Part 1 to the multi-agent platform we’ve built up to here.

More from the studio.

Back to blog