This is Part 6 of the series “Deep Dive: How Claude Code Works Under the Hood.” This post wraps up the architecture deep dive. Each post stands alone, but Parts 1-5 lay the groundwork.
Memory is one of those features that sounds simple until you try to build it. “Just remember things between sessions” --- easy to say, brutally hard to implement well for an AI agent. What counts as worth remembering? Where does it go? How does it get back into the model’s context without wasting tokens?
And then there is the system prompt --- the invisible foundation that shapes every response Claude Code gives you. Most people never think about it, but it is one of the most carefully engineered parts of the entire system.
I have been running Claude Code daily at ZenoLab for months, and understanding these two subsystems has fundamentally changed how I structure my projects and my interactions with the agent. Let me share what I have learned.
Memory: Durable Facts That Should Still Matter Next Session
The best definition of Claude Code’s memory system I have found is this: memory stores durable facts that should still matter the next time you open a session.
That sounds obvious, but the emphasis on “durable” and “still matter” is doing a lot of work. Most of what happens in a Claude Code session is transient --- file reads, command outputs, intermediate reasoning. None of that belongs in memory. Memory is for the things that transcend any single session.
The Four Memory Categories
Claude Code organizes memory into four categories, each serving a different purpose:
User memory stores facts about you as a developer. Your preferred coding style, your tech stack preferences, how you like commit messages formatted, which testing frameworks you favor. These are personal preferences that apply across all projects.
User memory examples:
- "Prefers functional React components over class components"
- "Uses Tailwind CSS for styling"
- "Likes concise commit messages, imperative mood"
- "Timezone: UTC+7"
Feedback memory captures corrections you have made. When you tell Claude “don’t use semicolons in TypeScript” or “always use named exports,” that correction gets stored so the agent does not repeat the mistake.
Feedback memory examples:
- "Do not add semicolons in TypeScript files (project uses no-semi rule)"
- "Use named exports instead of default exports"
- "Do not suggest installing packages without checking package.json first"
Project memory holds facts specific to the current repository. Architecture decisions, deployment conventions, naming patterns, API contracts. These are things any developer working on this project would need to know.
Project memory examples:
- "This project uses a monorepo with pnpm workspaces"
- "API routes follow /api/v1/{resource}/{id} pattern"
- "Database migrations run via prisma migrate"
- "CI/CD pipeline: GitHub Actions → Vercel"
Reference memory stores pointers to important files or documentation. Rather than memorizing the content of a file, the agent remembers where to find specific types of information.
Reference memory examples:
- "API types are defined in src/types/api.ts"
- "Environment variables documented in .env.example"
- "Design tokens in src/styles/tokens.ts"
What NOT to Store in Memory
This is just as important as what to store. Claude Code’s memory is not a dumping ground. Here is what does not belong:
File trees and directory structures. These change constantly. The agent can run ls or use the Glob tool to get the current structure whenever it needs it. Memorizing a file tree means memorizing something that will be stale by next session.
Code structure and implementation details. Do not store “the UserProfile component has three hooks and renders a form.” The agent can read the file. Store the decision (“we chose React Hook Form over Formik for forms”) not the current state.
Task status and to-do lists. Memory is not a task tracker. “Need to fix the login bug” is session context, not durable memory. If you need task tracking, use actual task management tools.
Secrets, tokens, and credentials. This should be obvious, but never store API keys, passwords, or tokens in memory. They end up in plain text files on disk. Use environment variables and secret managers.
Temporary debugging context. “The error was caused by a null pointer in line 47 of auth.ts” is not memory. It is session context. The error might be fixed by next session, and line 47 might be completely different code.
Memory vs Tasks vs CLAUDE.md
People often confuse these three systems, so let me clarify:
Memory is agent-managed. Claude Code decides what to store based on patterns in your interactions. You can view and edit it, but the agent is the primary author.
Tasks are user-managed. You tell Claude what to do, and it executes. Tasks are ephemeral --- they exist for one session or one interaction.
CLAUDE.md is developer-managed. You write it, you maintain it, and it provides persistent project context. It is a file in your repository, version-controlled and shared with your team.
The distinction matters because each system has different durability, authorship, and scope. Memory persists across sessions but is personal. CLAUDE.md persists across sessions and is shared. Tasks exist for a single session.
The System Prompt Assembly Pipeline
Now let me show you something that most people never see: how Claude Code’s system prompt is actually built.
The system prompt is not a single monolithic block of text. It is an assembly pipeline --- a sequence of builders, each pulling content from a specific source and adding it to the final prompt in a specific order.
The Build Order
The system prompt is assembled in this order:
- Identity --- Who is Claude Code? Base personality, capabilities, constraints.
- Tools --- What tools are available? Their schemas, descriptions, usage rules.
- Skills --- What skills are loaded? The lightweight index from Part 4.
- Memory --- Durable facts from the four memory categories.
- CLAUDE.md --- Project-specific instructions from CLAUDE.md files.
- Runtime context --- Current directory, git status, recent errors, session state.
┌─────────────────────────────────────┐
│ System Prompt │
├─────────────────────────────────────┤
│ 1. Identity (stable) │
│ "You are Claude Code..." │
├─────────────────────────────────────┤
│ 2. Tools (stable) │
│ Available tools and schemas │
├─────────────────────────────────────┤
│ 3. Skills (semi-stable) │
│ Skill index + loaded skills │
├─────────────────────────────────────┤
│ 4. Memory (semi-stable) │
│ User, Feedback, Project, Ref │
├─────────────────────────────────────┤
│ 5. CLAUDE.md (stable) │
│ Project instructions │
├─────────────────────────────────────┤
│ 6. Runtime (dynamic) │
│ cwd, git state, session info │
└─────────────────────────────────────┘
Each Builder Pulls from One Source Only
This is an underrated design principle: each builder in the pipeline has exactly one source of truth. The identity builder reads from the identity config. The tools builder reads from the tool registry. The memory builder reads from the memory store. No builder reaches into another builder’s domain.
Why does this matter? Traceability. When something goes wrong --- when the agent behaves unexpectedly or seems to have wrong instructions --- you can trace exactly which builder contributed which part of the prompt. There is no ambiguity about where a particular instruction came from.
As an indie developer at ZenoLab, I find this invaluable for debugging. If Claude Code is doing something I did not expect, I can check: Is it in my CLAUDE.md? Is it in memory? Is it a tool instruction? The pipeline structure tells me exactly where to look.
Stable vs Dynamic Content
Notice the annotations in the diagram above --- “stable” vs “semi-stable” vs “dynamic.” This separation is intentional.
Stable content (identity, tools, CLAUDE.md) rarely changes during a session. It gets assembled once at the start and stays fixed. This means the model can rely on these instructions being consistent throughout the conversation.
Semi-stable content (skills, memory) might change during a session. A new skill could be loaded, or a new memory could be added. But changes are infrequent --- maybe a few times per session at most.
Dynamic content (runtime context) changes constantly. The current directory, the git status, the list of recent files --- these are refreshed regularly to keep the agent grounded in the actual state of the workspace.
This hierarchy helps the compaction system (from Part 4) make better decisions. Stable content is never compacted. Semi-stable content is compacted conservatively. Dynamic content is compacted aggressively, since it can always be regenerated.
Memory Re-Injection: The Critical Feedback Loop
Here is an insight that took me a while to fully appreciate: if memory never re-enters model input, it is not actually guiding the agent.
Storing a memory is only half the system. The other half is re-injecting that memory into future conversations. Claude Code does this through the system prompt pipeline --- the memory builder pulls all relevant memories and includes them in every new session’s system prompt.
This means every conversation starts with your accumulated preferences, corrections, and project context already loaded. The agent does not have to re-learn that you prefer Tailwind over CSS modules, or that your project uses a specific deployment pattern. It knows from the first turn.
But this also means memory has a cost. Every stored memory consumes tokens in the system prompt. If you store hundreds of memories, that is hundreds of tokens that could be used for conversation context instead. This is why being selective about what goes into memory matters --- it is a budget, not a free resource.
Session starts
│
▼
Memory builder reads all stored memories
│
▼
Relevant memories injected into system prompt
│
▼
Agent starts with full context:
- Your preferences (User memory)
- Past corrections (Feedback memory)
- Project conventions (Project memory)
- Where to find things (Reference memory)
│
▼
First user message benefits from ALL prior learning
CLAUDE.md Layering
One of the most powerful features for team environments is CLAUDE.md layering. Claude Code does not just read one CLAUDE.md file --- it reads multiple, from different directory levels, and merges them.
The Layering Order
~/.claude/CLAUDE.md # User-level (your personal defaults)
↓
/project/CLAUDE.md # Project root (shared team conventions)
↓
/project/src/CLAUDE.md # Subdirectory (domain-specific rules)
↓
/project/src/api/CLAUDE.md # Deeper subdirectory (even more specific)
Each level can add or override instructions from the level above. This follows the same principle as CSS specificity or configuration cascading in tools like ESLint --- more specific rules take precedence.
Practical Example
At ZenoLab, my CLAUDE.md setup looks something like this:
~/.claude/CLAUDE.md (personal defaults):
- Use TypeScript strict mode
- Prefer functional programming patterns
- Write concise comments, avoid obvious ones
- Test files go next to source files, not in a separate directory
~/zenolab/CLAUDE.md (project root):
- This is an Astro project with React islands
- Styling: Tailwind CSS with custom design tokens in src/styles/
- Content lives in src/content/ using Astro content collections
- Deploy target: Vercel
- Node version: 20 LTS
~/zenolab/src/content/CLAUDE.md (content directory):
- Blog posts use YAML frontmatter: title, description, date, tags, draft
- Date format: YYYY-MM-DD
- Tags are lowercase, hyphenated
- Always include a meta description for SEO
When Claude Code works on a blog post, it has all three levels of context. It knows my personal preferences (from user-level), the project architecture (from project root), and the content conventions (from the content directory). No single file has to contain everything.
Why Layering Beats a Single File
I tried the single CLAUDE.md approach first. It worked for small projects. But as the project grew and I started working in different areas of the codebase, a single file became bloated and contradictory. Frontend instructions mixed with API instructions mixed with content instructions.
Layering lets each directory own its context. The API directory knows about REST conventions. The content directory knows about frontmatter schemas. The component directory knows about React patterns. And none of them need to know about the others.
This also works well for teams. The project-root CLAUDE.md captures shared conventions. Individual developers add their personal CLAUDE.md for preferences that do not need to be shared. Subdirectories add domain-specific rules that only matter when working in that area.
Bringing It All Together
Memory, system prompts, and CLAUDE.md form a three-part context system:
- Memory provides learned, agent-curated context that evolves over time
- System prompts provide the foundational identity and capability context
- CLAUDE.md provides developer-authored, version-controlled project context
Each serves a different need, but they all flow through the same pipeline and end up in the same place: the model’s input context at the start of every turn. Together, they give Claude Code the ability to be a genuinely contextual assistant that knows your preferences, your project, and your patterns.
For an indie developer like me, this means I can start a fresh session and Claude Code already knows how I work, what my project looks like, and what conventions to follow. There is no “warm-up” period. The agent is productive from turn one.
What’s Next
This wraps up the core architecture of Claude Code. Over six posts, we have covered the agentic loop, tool execution, subagent orchestration, skill loading, context compaction, permissions, hooks, memory, and system prompt assembly.
If you are building your own tools on top of Claude Code, or just want to use it more effectively, understanding these internals gives you a massive advantage. You know why certain things work the way they do, and you can structure your projects to work with the system instead of against it.
Coming up next in the series, we will shift from architecture to practice --- looking at real-world workflows, advanced configurations, and patterns I have developed running Claude Code in production at ZenoLab. Stay tuned.