Prompt InsightsOpen Prompt Builder

Glossary // Context compaction

What Is Context Compaction? How AI Agents Summarize Their Own Context (2026)

Updated

Context compaction is what an AI agent does when its context window is nearly full: it replaces the older part of the conversation with a model-written summary and continues from that summary. The session survives, the token count drops, and the agent keeps working. What it loses is everything the summary did not keep, and the agent has no way to know what that was.

Every model has a fixed context window, the maximum number of tokens it can read at once. A long coding or research session fills it with file contents, tool output, error logs, and back-and-forth. Something has to give, and the three options are to stop, to drop the oldest turns (truncation), or to compress them (compaction). Agent harnesses have mostly settled on compaction, because truncation throws away the start of the task, which is usually where the instructions were.

How compaction works in practice

The mechanics are similar across tools. When the context crosses a threshold, the harness sends the conversation so far to the model with an instruction to summarize it for continuation: the goal, the work done, the current state, the next steps. The summary replaces the old turns, the most recent turns are often kept verbatim, and the session continues.

  • Claude Code compacts automatically as the window fills and exposes a /compact command to do it on demand. Text after the command steers the summary, for example /compact keep the list of failing tests and why we chose the new schema. /clear is the opposite move: drop the context entirely and start fresh.
  • Other agent harnesses follow the same pattern with their own thresholds and summary prompts; Docker's agent tooling, for instance, documents both automatic and on-demand compaction plus trimming of large tool results.
  • API-level context management offers a lighter alternative: instead of summarizing everything, the harness clears stale tool results (a file listing from forty steps ago) and keeps the conversation itself intact.

What compaction drops

A summary is a bet about what will matter later, made before anyone knows. The things that lose that bet are predictable:

  1. Constraints stated once. "Do not touch the billing module" said at turn 3 is exactly the kind of sentence a summary compresses to "working on refactor". This is sometimes called governance decay: the rules fade because they are old, not because they stopped applying.
  2. Negative results. What was tried and failed rarely makes the summary, so the agent tries it again.
  3. Exact strings. Error messages, file paths, version numbers, and IDs get paraphrased, and a paraphrased error is useless for searching.
  4. The reason behind a decision. The summary keeps "switched to approach B" and drops why approach A was ruled out, which invites a switch back.
  5. User preferences. Style and format corrections given mid-session are the first thing to go.

Practitioners have been vocal about this. Our coverage of Claude's context compaction problem traces a wave of complaints to post-compaction behaviour in long agentic sessions, and to community tools built to patch it. The lossless memory approach goes further and argues memory should never be summarized at write time at all, only filtered at read time.

How to make a session survive compaction

The fix is not to avoid compaction but to stop relying on the conversation as the only record.

  • Put durable rules outside the conversation. Project instruction files (CLAUDE.md, AGENTS.md, or your harness's equivalent) are loaded as standing context rather than as chat turns, so they do not depend on surviving a summary. Constraints that apply to the whole project belong there, not in turn 3.
  • Keep a working notes file. Ask the agent to maintain a short NOTES.md or progress.md with the goal, decisions with reasons, dead ends, and next steps, and to update it before long operations. After compaction, the first instruction is to reread it.
  • Steer the summary. When you compact by hand, say what to keep. A one-line instruction beats a perfect default prompt.
  • Split tasks instead of stretching sessions. A session that needs three compactions is usually two or three tasks. Finish one, write the state down, and start the next one clean.
  • Check the first action after a compaction. It is the most likely point for the agent to redo a failed approach or break a rule, which makes it a natural place for a trace check of the kind described in our silent agent failures guide.

A compaction prompt for your own harness

If you build your own agent loop, the summary prompt is yours to write. This one is structured so the things that usually get lost have a slot:

Summarize this session so another instance can continue the work without the original transcript.
Use exactly these sections:
GOAL: the task in one or two sentences, in the user's words where possible.
CONSTRAINTS: every rule or preference the user stated, verbatim, even if it was said once.
DECISIONS: each decision made, with the reason in one line.
DEAD ENDS: approaches tried and why they failed, including exact error messages.
STATE: files changed, commands that matter, current status of tests or outputs.
NEXT: the next three concrete steps.
Do not drop a constraint to save space. Shorten anything else first.

The last line encodes the priority: constraints and dead ends are cheap in tokens and expensive to lose.

Compaction, memory, and trust

Compaction is housekeeping inside one session; agent memory is what an agent deliberately keeps across sessions. They meet at the summary: whatever compaction writes becomes the agent's belief about its own past. That makes the summary a place where a prompt injection read earlier in the session can be laundered into "the user asked for this". If your agent read untrusted content before compacting, treat the summary with the same suspicion as the content.

Context compaction in the news

READY TO ASCEND

Get prompt-engineering insight by email

New terms, prompt packs, and the news that changes how you prompt. No noise.

Questions

What does context compaction mean?

It means summarizing the earlier part of an AI conversation or agent session so it takes fewer tokens, then continuing with that summary in place of the original turns. The agent keeps working, but from a compressed account of what happened rather than the full record.

What is /compact in Claude Code?

It is the command that triggers compaction by hand. Claude Code also compacts automatically as the context window fills. You can pass instructions after the command, such as which decisions, files, or failing tests to keep, and the summary will prioritise them.

Is compaction the same as memory?

No. Compaction is lossy housekeeping inside one session. Memory is information deliberately written somewhere outside the context window, a file, database, or index, so it can be read back later, including in a new session. Good agents use memory to protect what compaction would otherwise drop.

Why does my agent get worse after compacting?

Because the summary dropped something the work depended on: a constraint you stated once, a decision and its reason, a path that was already tried and failed, or the exact error text. The agent is not broken, it is working from a shorter story than the one you remember telling it.

Should I compact or start a new session?

Start fresh when the next task is different from the last one; carry the important state across in a short notes file. Compact when you are mid-task and need the thread of the current work, and give the compaction explicit instructions about what to keep.

Related terms

Prompt packs for this