This is Part 8 of the series “Deep Dive: How Claude Code Works Under the Hood.” Last time we covered error recovery — how Claude Code classifies failures and routes them to targeted recovery strategies. Now we step into one of the most architecturally interesting features: persistent task management.
Beyond Flat Checklists
Early in my work at ZenoLab, I built task tracking into an agent the obvious way: a JSON array of strings. Check items off as they complete. Simple, right?
It fell apart the moment tasks had dependencies. “Run database migration” has to finish before “seed test data.” “Build frontend” and “build backend” can run in parallel, but “deploy” needs both. A flat list can’t express this. You end up with fragile ordering logic and race conditions.
Claude Code solves this with persistent task DAGs — directed acyclic graphs where each task knows its dependencies, its status, and its place in the execution order. Let me show you how it works.
Tasks as JSON Files
The first design decision that matters: tasks are stored as JSON files on disk, not in memory.
{
"id": "task_a1b2c3",
"title": "Refactor authentication module",
"status": "in_progress",
"dependencies": ["task_x9y8z7"],
"created_at": "2026-04-22T10:30:00Z",
"updated_at": "2026-04-22T11:15:00Z",
"result": null
}
This is a deliberate choice that pays off in two critical scenarios:
-
Context compression: When the conversation gets compacted (as we discussed in the error recovery post), task state isn’t lost. The JSON files survive independently of the conversation history. After compaction, the agent can re-read the task files and pick up exactly where it left off.
-
Session restarts: If you close your terminal and come back tomorrow, the tasks are still there. They’re files, not ephemeral state. This is what makes Claude Code feel like it has a persistent memory for work-in-progress.
The .tasks/ directory becomes the source of truth. The conversation is just the interface to it.
Dependency Resolution
Each task has a dependencies array listing the IDs of tasks that must complete before it can start. The resolution logic is straightforward but powerful:
function getReadyTasks(tasks) {
return tasks.filter(task => {
if (task.status !== 'pending') return false;
return task.dependencies.every(depId => {
const dep = tasks.find(t => t.id === depId);
return dep && dep.status === 'completed';
});
});
}
A task is “ready” when all its dependencies are completed and it’s still in pending status. This creates a natural execution order without anyone having to manually sequence things.
Here’s a concrete example. Say you’re working on a feature that requires:
- Update the database schema
- Generate new TypeScript types (depends on 1)
- Update API endpoints (depends on 2)
- Update frontend components (depends on 2)
- Write integration tests (depends on 3 and 4)
- Update documentation (depends on 3 and 4)
Tasks 3 and 4 can run in parallel once task 2 finishes. Tasks 5 and 6 can run in parallel once both 3 and 4 finish. The DAG captures this naturally — no manual scheduling needed.
[Schema] → [Types] → [API Endpoints] ──→ [Integration Tests]
↘ ↗
[Frontend] ────────→ [Documentation]
When the agent completes a task, it checks if any downstream tasks are now unblocked. If they are, they move into the ready queue. This is the “completed tasks unblock dependents” principle in action.
Status Transitions
Every task follows a simple state machine:
pending → in_progress → completed
The transitions are:
- pending: Task exists but isn’t ready to run (dependencies incomplete) or hasn’t been picked up yet
- in_progress: Task has been claimed and work is underway
- completed: Task finished successfully with a result
function transitionTask(task, newStatus, result = null) {
const validTransitions = {
'pending': ['in_progress'],
'in_progress': ['completed', 'pending'], // can revert if interrupted
};
if (!validTransitions[task.status]?.includes(newStatus)) {
throw new InvalidTransitionError(task.status, newStatus);
}
task.status = newStatus;
task.updated_at = new Date().toISOString();
if (result) task.result = result;
writeTaskFile(task);
checkDependents(task); // unblock waiting tasks
}
Notice that in_progress can transition back to pending. This handles the case where a task gets interrupted (session crash, context overflow recovery) and needs to be re-attempted. Without this, interrupted tasks would be stuck in in_progress forever — a classic issue in queue-based systems.
Background Tasks: Daemon Threads for Slow Operations
Some operations are slow. npm install can take 30 seconds. A full pytest suite might run for minutes. Waiting synchronously for these would freeze the agent — it can’t think about anything else while a subprocess is running.
Claude Code handles this with background tasks: daemon-like threads that run slow operations while the main agent loop continues thinking.
function spawnBackgroundTask(command, taskId) {
const process = spawn(command, {
detached: true,
stdio: ['ignore', 'pipe', 'pipe'],
});
backgroundTasks.set(taskId, {
process,
stdout: [],
stderr: [],
status: 'running',
startedAt: Date.now(),
});
process.stdout.on('data', (data) => {
backgroundTasks.get(taskId).stdout.push(data.toString());
});
process.on('exit', (code) => {
const task = backgroundTasks.get(taskId);
task.status = code === 0 ? 'completed' : 'failed';
task.exitCode = code;
});
}
The background process runs independently. Its output (stdout, stderr) is captured into a buffer. The main agent loop can check on it periodically or be notified when it finishes.
This is the kind of thing you don’t appreciate until you’re five minutes into a large monorepo install and realize the agent is still responsive, still answering questions, still working on other tasks while node_modules populates.
The Drain-Before-Call Pattern
Here’s where it gets elegant. Before the agent sends a new prompt to the model (before each “think” step), it drains the results from any completed background tasks.
async function agentLoop() {
while (running) {
// Drain completed background tasks
const completedResults = drainCompletedBackgroundTasks();
if (completedResults.length > 0) {
injectResults(conversation, completedResults);
}
// Now ask the model to think
const response = await callModel(conversation);
// ... process response, execute tools, etc.
}
}
The “drain-before-call” pattern means the model always has the latest information when it starts thinking. If npm install finished while the model was writing code, the results are there in the next turn. If pytest found failures, those are injected before the model plans its next move.
This is much better than interrupting the model mid-response with async notifications. The model gets a clean, complete picture of the world state at the start of each thinking cycle.
Cron Scheduler: Store Future Intent Now
Background tasks handle “do this slow thing now.” But what about “do this thing every hour” or “check for new issues at 9 AM”? That’s where the cron scheduler comes in.
The design philosophy is simple: store future intent now, trigger it later.
{
"id": "schedule_r4s5t6",
"description": "Check for stale pull requests",
"cron": "0 9 * * 1-5",
"command": "review_stale_prs",
"lastFired": "2026-04-21T09:00:00Z",
"nextFire": "2026-04-22T09:00:00Z",
"durability": "durable",
"enabled": true
}
A schedule record contains:
- cron expression: Standard cron syntax for when to fire (
0 9 * * 1-5= 9 AM on weekdays) - command: What to execute when it fires
- lastFired / nextFire: Timestamps for preventing duplicate fires
- durability: Whether this schedule survives session restarts
Duplicate-Fire Prevention
The lastFired timestamp solves a subtle but important problem. If the agent restarts at 9:01 AM, should the 9 AM job fire again? Without tracking, it would — the scheduler sees “it’s past 9 AM and the cron matches” and fires.
With lastFired, the check becomes:
function shouldFire(schedule, now) {
const nextFire = getNextCronTime(schedule.cron, schedule.lastFired);
return now >= nextFire;
}
It calculates the next fire time based on the last fire time. If the job already fired at 9:00 AM, the next fire is tomorrow at 9:00 AM. No duplicate execution.
Durable vs Session-Only Schedules
Not all schedules should persist. Some make sense only within a specific work session:
- Durable schedules survive restarts. “Check for stale PRs every weekday morning” is a standing instruction that should persist indefinitely. These are written to disk.
- Session-only schedules disappear when the session ends. “Re-run tests every 10 minutes while I’m refactoring” is temporary. These live in memory only.
This distinction keeps the schedule store clean. Without it, you’d accumulate dozens of one-off schedules that no longer make sense.
Unified Notification Model
Background tasks and scheduled tasks share a common notification model. When something completes — whether it was a background npm install or a scheduled review_stale_prs — the result flows through the same channel:
function notify(event) {
// Same structure regardless of source
const notification = {
type: event.source, // 'background_task' | 'scheduled_task' | 'user_action'
taskId: event.taskId,
status: event.status,
result: event.result,
timestamp: Date.now(),
};
notificationQueue.push(notification);
}
The drain-before-call pattern we saw earlier pulls from this unified queue. The model doesn’t need to know whether a result came from a background task, a cron job, or a user action. It’s all just “new information since last time I thought.”
This is a pattern I’ve adopted directly in my ZenoLab projects. Having a single notification path instead of separate handlers for each source eliminates an entire class of “but what about this case?” bugs.
Putting It All Together
Let me trace through a realistic scenario. You tell Claude Code: “Set up the database, generate types, update the API, and write tests.”
- The agent creates four tasks with the right dependency structure
- Task 1 (database migration) starts immediately as a background task
- The agent continues the conversation with you while the migration runs
- On the next think cycle, the drain step picks up the completed migration
- Tasks 2 (generate types) is now unblocked and starts
- Types complete, tasks 3 (API) and 4 (tests — which depends on 3) are evaluated
- Task 3 starts, task 4 waits
- Task 3 completes, task 4 is unblocked and starts
- All tasks complete, the agent reports the full result
At no point did the agent stop to wait. At no point did it lose track of what was done and what remained. The task DAG handled the sequencing, the background execution handled the parallelism, and the drain pattern kept the model informed.
What’s Next
In Part 9, we enter genuinely new territory: agent teams. Not a single agent doing tasks sequentially, but multiple AI agents working together as persistent teammates with their own inboxes, lifecycles, and coordination protocols. This is where Claude Code starts to feel less like a tool and more like a team.