This is Part 4 of the series “Deep Dive: How Claude Code Works Under the Hood.” If you missed the earlier posts, start from Part 1 for the full picture. Each post stands on its own, but they build on each other.
If you have ever watched your context window fill up and wondered why Claude Code suddenly got less sharp mid-session, this post is for you. I have spent a lot of time studying how Claude Code handles two related problems: loading the right knowledge at the right time, and keeping the conversation window from overflowing with stale data.
These two subsystems --- skills and context compaction --- are what separate a prototype AI agent from one that can actually sustain long work sessions. Let me walk you through how they work.
The Two-Layer Skill Loading Pattern
Most people assume that Claude Code loads all of its capabilities into memory at the start of a session. That would be the simple approach. It would also be a terrible one.
Here is the actual pattern: advertise cheaply, load on demand.
At startup, Claude Code does not inject full skill definitions into the system prompt. Instead, it injects a lightweight index --- just the skill names, a one-line description, and a trigger hint. The full skill content stays on disk until the agent actually needs it.
The Cookbook Analogy
Think of it like a cookbook shelf in your kitchen. You know which shelf holds the Italian cookbook and which has the baking guide. You do not open every cookbook and spread all the pages across your counter before you start cooking. You glance at the spines, pick the one you need, and open it to the relevant page.
That is exactly what Claude Code does with skills. The “spines” are cheap metadata lines in the system prompt. The full content is a SKILL.md file that only gets loaded when the agent decides it is relevant to the current task.
Why This Matters: The 20k Token Problem
I learned this lesson the hard way while building tools at ZenoLab. Early in my experiments with agent architectures, I tried the naive approach: dump everything into the system prompt. Instructions for code review, deployment, testing, documentation, formatting, security checks --- all of it, all the time.
The result? About 20,000 tokens of context burned before the user even typed their first message. That is roughly 15% of a standard context window gone before any work happens. Every subsequent turn has less room for actual reasoning, code, and conversation history.
The two-layer pattern solves this. The index costs maybe 200-400 tokens total. Each skill, when loaded, costs only what that specific skill needs. If you never ask about deployment, the deployment skill never enters the window.
# What the agent sees at startup (cheap index)
Available skills:
- code-reviewer: Reviews code for style and best practices
- deploy-helper: Manages deployment workflows
- test-writer: Generates test suites from source code
- doc-generator: Creates documentation from code
# What gets loaded on demand (full content)
[Only loaded when the agent calls ToolSearch or Skill tool]
SKILL.md Files and the Metadata Contract
Each skill lives in a SKILL.md file with YAML frontmatter at the top. The frontmatter is the metadata contract --- it tells the loading system what the skill is, when to suggest it, and what it needs.
---
id: code-reviewer
name: Code Reviewer
description: Reviews code files for style, performance, and best practices
triggers:
- pattern: "review"
description: "Triggered when user asks for code review"
---
# Code Reviewer Skill
When reviewing code, follow these steps:
1. Check for obvious bugs and logic errors
2. Evaluate naming conventions and readability
3. Look for performance concerns
...
The frontmatter fields serve the index layer. The body serves the execution layer. They are cleanly separated, and that separation is what makes lazy loading possible.
Model Autonomy in Skill Loading
Here is something that surprised me when I first dug into this: the model itself decides when to load a skill. There is no hard-coded rule that says “if the user mentions ‘review,’ load the code-reviewer skill.” Instead, the agent sees the available skill list, evaluates the current request, and makes a judgment call about whether loading a particular skill would help.
This is genuine agent autonomy. The system trusts the model to manage its own knowledge retrieval. It is the same principle behind tool use --- the model decides which tools to call and when. Skills extend that same pattern to knowledge management.
Context Compaction: The Four-Lever Strategy
Now we get to the second half of the puzzle. Even with lazy skill loading, long sessions accumulate context. You run twenty tool calls, read ten files, generate a bunch of code. The conversation history grows, and eventually you start bumping against the window limit.
Claude Code handles this with what I think of as a four-lever compaction strategy. Each lever is progressively more aggressive, and the system escalates through them as pressure on the context window increases.
Lever 0: Persisted Output
This is the gentlest form of compaction, and it happens before you even notice pressure. When a tool produces a large result --- say, reading a 500-line file or running a command that dumps pages of output --- the system can persist that output to disk and replace the in-context version with a reference.
# Before compaction
[Tool Result: Read file /src/components/App.tsx]
(500 lines of React component code)
# After Lever 0
[Tool Result: Read file /src/components/App.tsx]
(Output saved to disk. Key findings: React component with 3 hooks,
renders a dashboard layout. Full content available if needed.)
The detail is not gone. It has been relocated. The agent can re-read the file if it needs the exact content again. But for most follow-up reasoning, the summary is sufficient.
Lever 1: Micro-Compact
This lever targets old tool results specifically. As the conversation progresses, earlier tool outputs become less relevant to the current task. Micro-compaction replaces those old results with compact placeholders.
Think of it like this: you do not need the full output of a git status you ran twenty turns ago. You needed it then. Now, a placeholder that says “checked git status, working tree was clean” is enough context to keep the conversation coherent.
# Turn 5 (original)
[Tool: Bash] git status
On branch main
Your branch is up to date with 'origin/main'.
Changes not staged for commit:
(use "git add <file>..." to update what will be committed)
modified: src/index.ts
modified: src/utils/helpers.ts
# Turn 25 (micro-compacted)
[Tool: Bash] git status → 2 modified files (src/index.ts, src/utils/helpers.ts)
The key property of micro-compaction: it is lossless for the current task. The agent can always re-run the command if it needs fresh data.
Lever 2: Auto-Compact
This is where the LLM itself gets involved in compaction. When the context window fills past a threshold, the system triggers an automatic summarization pass. The model reads the full conversation history and produces a condensed version that preserves:
- Key decisions made
- Current task state
- Important findings
- File modifications performed
- Outstanding questions or blockers
This is more aggressive than micro-compaction because it restructures the entire conversation, not just individual tool results. But it is still automatic --- the user does not need to do anything.
I have seen auto-compact reduce a 90k-token conversation down to 15k tokens while retaining all the information needed to continue working. That is an 80%+ compression ratio without losing the thread of the work.
Lever 3: Manual Compact
Sometimes the agent itself recognizes that context is getting bloated and explicitly triggers a compaction. This is the most aggressive lever --- the agent essentially says “let me summarize where we are so we can keep working efficiently.”
You can also trigger this yourself with the /compact command in Claude Code. It forces a full summarization pass, which is useful when you have been exploring a codebase and want to pivot to a different task without losing the context of what you learned.
You: /compact
Claude: I'll summarize our conversation so far...
We've been investigating a performance issue in the dashboard component.
Key findings:
- The re-render is caused by an unstable reference in useEffect deps
- The fix involves memoizing the config object with useMemo
- We've already updated src/hooks/useDashboard.ts
- Tests are passing
Ready to continue. What's next?
The Core Insight: Relocation, Not Deletion
Here is the mental model that ties all of this together: compaction is relocating detail, not deleting history.
Nothing is permanently lost during compaction. Lever 0 moves data to disk. Lever 1 replaces verbose outputs with summaries. Lever 2 restructures the conversation into a denser format. Lever 3 does the same thing but on explicit request.
At every level, the agent retains the ability to recover detail. It can re-read files, re-run commands, or ask you to clarify something from earlier. The compacted context gives it enough to know what happened and where to look if it needs more.
This is fundamentally different from truncation, which is what simpler systems do. Truncation just chops off the oldest messages. That is like ripping pages out of a notebook --- you lose whatever was on those pages, and if it was important, too bad.
Compaction is more like rewriting your notes in a smaller notebook. The information density goes up, the volume goes down, and you can always go back to the original sources if the summary is not enough.
Practical Implications for Long Sessions
Understanding these mechanics has changed how I work with Claude Code at ZenoLab. A few practical takeaways:
Structure your work in phases. If you are doing exploration followed by implementation, use /compact between phases. The exploration context gets summarized, and you start implementation with a clean, dense summary instead of pages of file listings.
Trust the lazy loading. Do not try to front-load context by pasting documentation into the conversation. Let the skill system do its job. The agent will pull in what it needs.
Watch for compaction artifacts. Occasionally, auto-compact will summarize away a detail you still need. If the agent seems to have “forgotten” something, it probably got compacted. Just re-state the detail or point it back to the relevant file.
Use SKILL.md files for your own projects. If you have complex workflows that require specific instructions, encode them as skills rather than pasting them into every conversation. This is the single biggest context-saving move you can make.
What’s Next
In Part 5, we will dig into the permissions and hooks system --- the safety architecture that controls what Claude Code can and cannot do. This is where things get interesting from a security perspective, especially if you are running Claude Code in auto mode or in CI pipelines. We will cover the four-stage permission pipeline, pattern matching rules, and how hooks let you observe and control tool execution without breaking the agent’s control flow.