---
title: Tokenmaxxing
description: Pruning your system prompt won't save you tokens; understanding
  what actually happens when you interact with an LLM will.
date: 2026-09-06
tags:
  - ai
  - agents
url: https://skurekjakub.dev/blog/tokenmaxxing
author: Jakub Skurek
---

> Agent instructions:
> Take page paths from the index rather than constructing them — guessed paths 404. Links inside this document already point at `.md`; following them, or sending `Accept: text/markdown`, keeps you in Markdown.
> When citing this page, use the `url` in its frontmatter — the HTML page; this `.md` address is a rendering of it.
>
> - [llms.txt](https://skurekjakub.dev/llms.txt) — every page with a Markdown twin, with descriptions.
> - [skill.md](https://skurekjakub.dev/skill.md) — how to read this site as an agent.

Over the last two months, Claude Code sent 1.18 billion tokens to the Anthropic API when working on this website alone, and received 6.26 million back. That is a ratio of roughly 200 to 1. In other words, for every token the model produced, it read nearly two hundred.

Nothing about these numbers should look surprising to experienced users. Still, it's understandable that for some, seeing a billion tokens on their `/usage` makes them want to optimise everything about their setup. Shave ten lines off `AGENTS.md`, prune the system prompt, or disable MCP servers and plugins to save a few hundred tokens. This need to "tokenmaxx" one's setup is something I also went through when delving deeper into the AI rabbit hole for the first time.

While there is nothing inherently incorrect about this approach, and it does do exactly what it says on the tin, I want to use this post to prove that all these efforts get dwarfed (on the scale of orders of magnitude) by simply understanding how harness (client) <-> LLM provider communication works and applying proper hygiene to session management.

## Context composition

To get an intuition for what an average conversation with an LLM contains, let's start by looking at context windows composition. There is a tendency to assume the system prompt, tool definitions, or harness instructions eat up much of a model's context. However, looking at transcripts from actual sessions proves otherwise:

**Interactive chart.** What the window was made of across 79 sessions on 2026-09-04, in tokens: tool output is 42.9% of the main thread, led by bash output at 23.1%; the session prefix the transcript never stores is 17.8%.

| Main session      | Share | Largest kind inside                               |
| :---------------- | ----: | :------------------------------------------------ |
| Tool output       | 42.9% | Bash output 23.1%                                 |
| Session prefix    | 17.8% | System prompt, tools, skills, memory, hooks 17.8% |
| Agent writes      | 16.6% | Write calls 10.1%                                 |
| User prompts      | 10.3% | User prompts 10.3%                                |
| Tool invocation   |  7.4% | Bash calls 5.8%                                   |
| Agent replies     |  3.5% | The agent's replies 3.5%                          |
| Harness reminders |  1.3% | System reminders and hooks 1.3%                   |

| Subagents         | Share | Largest kind inside                               |
| :---------------- | ----: | :------------------------------------------------ |
| Tool output       | 54.5% | Bash output 27.1%                                 |
| Session prefix    | 20.4% | System prompt, tools, skills, memory, hooks 20.4% |
| Agent replies     |  6.4% | The agent's replies 6.4%                          |
| Harness reminders |  5.7% | System reminders and hooks 5.7%                   |
| Agent writes      |  5.0% | Write calls 4.4%                                  |
| Tool invocation   |  5.0% | Bash calls 3.4%                                   |
| User prompts      |  2.7% | User prompts 2.7%                                 |

_Context window composition across 79 coding sessions._

Across these sessions, the content in the context window is rougly split into:

- **Tool output (42.9%)** — shell commands and common tool calls (search, read, write, edit) make up the majority of each context.
- **Session prefix (17.8%)** — initial session bootstrap. Contains the system prompt, schemas for mandatory tools and subagents, MCP server definitions, among others.
- **File writes (16.6%)** — file edits made by the agent persist in context for the rest of the session.
- **User prompts (10.3%)** — task descriptions, follow-up answers, and manual steering.
- **Tool invocations (7.4%)** — agent tool invocations. Contains the schema, props, and parameters.
- **Agent responses (3.5%)** — the model's conversational replies and reasoning text (raw, opaque, or as a summary, varies per harness).
- **Harness reminders (1.3%)** — `<system-reminder>` blocks and hook output injected mid-session. Rule reminders, file-changed notices, skill listings after a compaction, custom reminders.

Well over half of the context window is raw tool output and code the agent generated itself. The next highest chunk is called the _session prefix_. It is unavoidable, injected by the harness itself, and serves to ground the model and the conversation. More on the prefix later as this is where most of the common token optimization advice concetrates. The rest of the window is a mixed bag of toolcalls, user prompts and various minutia that varies depending on session goals.

> **Tip: Inspect your own session footprint**
>
> Run `/context` inside an active Claude Code session to see a breakdown of loaded tools, system prompts, and message history. Alternatively, have an agent inspect the raw JSONL transcripts under `~/.claude/projects/` to see what dominates your sessions. You can do the same with other harnesses like Copilot or Codex.

## The context window

Next is the context window. An LLM has no memory between turns. Every time you send a prompt or a tool finishes running, the client packages the entire conversation history and sends it back to the API as a single JSON file.

Because the session is append-only, every turn re-reads the entire history from the top. Here is what that looks like across one long research session on this repo:

**Interactive chart.** Context per call in the "this post's research" session: 249 calls, a peak of 451k tokens, 3 compactions, 51.81M tokens sent over its life. Across all sessions the window just before a compaction was 275k tokens at the median.

| Session                      | Peak | Calls | Compactions | Sent over the session |
| :--------------------------- | ---: | ----: | ----------: | --------------------: |
| attention lab port           | 591k |   223 |           2 |                56.19M |
| markdown routing post        | 529k |   425 |           1 |               109.73M |
| this post's research         | 451k |   249 |           3 |                51.81M |
| one long read, no compaction | 417k |    49 |           0 |                14.81M |

| Compactions in "this post's research" | Window before | Window after |
| :------------------------------------ | ------------: | -----------: |
| call 49                               |          277k |          61k |
| call 153                              |          451k |          65k |
| call 187                              |          206k |          70k |

_Context growth per API call in a single session, with compactions marked._

The context starts around 48k tokens and climbs steadily as the agent searches code, inspects files, and runs commands. Eventually, as the context window approaches capacity, the harness triggers **compaction**.

Compaction replaces the accumulated conversation history with a model-generated summary. Claude Code runs this automatically as the window fills up, or you can trigger it manually with `/compact`.

> **Tip: Keep an external plan checkpoint**
>
> Before asking an agent to run a large refactor in an already long session, ask it to write its plan and checklist to a file on disk. The agent then still has a stable reference point even if the compaction erases important details from its context.

## Prompt caching

Re-reading hundreds of thousands of tokens on every prompt and toll call sounds expensive. If you start from a 300 thousand token session and reread it twenty times, that adds up to 6 million input tokens. With most vendors pricing even baseline models like Sonnet at $2 per 1 million _input_ tokens, how is everyone not bankrupt already?

The secret is **prompt caching** (you might also see it called the KV cache).

When an API request comes in, the provider checks for an exact prefix match against prior requests. If the prefix matches, the model reuses its cached values computed for the session so far in prior turns.

Vendors price this explicitly. On Anthropic's API:

- A **cache read** costs 0.1× (10%) of the base input token price.
- A **cache write** into the 5-minute cache costs 1.25× the base price.
- A **cache write** into the 1-hour cache costs 2.0× the base price.

Taken from their documentation (other vendors use identical pricing structures):

![API pricing pulled from Anthropic's documentation September 5, 2026](https://skurekjakub.dev/blog/tokenmaxxing/pricing.png)

_API pricing excerpt pulled from Anthropic's documentation September 5, 2026_

Claude Code, for example, uses the 1-hour cache lifetime so pauses between prompts do not immediately drop the cache.

Here is a breakdown for the sessions that worked over this website. You can immediately see that the **majority** of token reads were cached.

**Interactive chart.** Of 1179.63M billed input tokens, 98.0% were served from a warm cache and 2.0% were computed cold, as cache writes or uncached.

| Cache                              |   Tokens | Of billed input |
| :--------------------------------- | -------: | --------------: |
| Warm: read from cache              | 1156.15M |           98.0% |
| Cold: written to cache or uncached |   23.48M |            2.0% |

| Whole-window re-write, by cause            | Calls | Written |
| :----------------------------------------- | ----: | ------: |
| first call of a session                    |    53 |   2.28M |
| just after a compaction                    |    25 |   1.16M |
| more than an hour since the last call      |    15 |   2.76M |
| more than five minutes since the last call |     5 |    667k |
| something earlier in the window changed    |    20 |   3.11M |

_Billed input tokens served warm from the cache versus computed cold, plus triggers for major cache rewrites._

Across the roughly 1.2 billion input tokens consumed by this project:

- **98.0% were cache reads**.
- **2.0% were 1-hour cache writes**. A large portion of the writes is expected — each fresh agent turn needs to be processed, cached and appended to the session.

Because cache reads cost a tenth of the standard rate (a fortieth on Fable 5.1), they accounted for 69% of the input rate consumption, and the 2% of tokens written to the cache consumed the rest. Claude Code prices each session locally at API list rates, which is the figure `/usage` shows as the session total. Summed over every session in the corpus, per model at that model's rates, it comes to:

|                                          | At API list price |
| ---------------------------------------- | ----------------- |
| Input, as billed                         | $1,121            |
| Input, had nothing been cached           | $8,737            |
| Input, had every token been a cache read | $792              |
| Output                                   | $254              |

The effective input price came to **0.13×** the base rate. Prompt caching saved a tidy 87% on the un-cached input cost without any special setup, and the subagents those sessions spawned added another $173 of input on the same basis. These numbers use API pricing. The sessions themselves ran on a Max subscription, for which Anthropic publishes no per-token rate (but is clearly infinitely cheaper...and likely massively subsidized).

## What disrupts prompt caching

Since we now know that prompt cache is the golden goose protecting us common folk from turning destitue, we clearly need to look into what breaks it and how to prevent it.

The prompt cache is strictly an ordered prefix cache, ordered as such:

```text
[Tool Definitions] → [System Prompt & Instructions] → [Message History]
```

Ordered means that if a single character changes anywhere in that chain, the cache breaks from that point forward. All such content is counted as a fresh read and billed at full API pricing, or consumes a much larger chunk from allocated limits in flat subscriptions.

The following widget simulates various actions that impact the shape of the prompt and its cache. Selecting different scenarios highlights what section of the cache needs a re-read as a consequence:

**Interactive chart.** The cache is a prefix in the order tools, system, messages, and a change at one level re-writes that level and everything after it. On a window of 230k tokens a plain next turn costs 31k base-token equivalents; the table prices the alternatives.

| What happens            | Read | Written | Costs | Explanation                                                                                                                                       |
| :---------------------- | ---: | ------: | ----: | :------------------------------------------------------------------------------------------------------------------------------------------------ |
| Send the next turn      | 226k |    4.0k |   31k | Everything before the new turn is read back from the cache; only the new turn is written.                                                         |
| Disable an MCP server   |    0 |    230k |  460k | The tool definitions change, and they are the first block, so every block after them is written again.                                            |
| Change the effort level |  46k |    184k |  373k | The thinking and effort configuration is rendered into the prompt, so the whole scroll is written again; the tools and the system prompt survive. |
| Switch fast mode on     |  28k |    202k |  407k | The speed setting is rendered into the system prompt, so it invalidates the system and message caches.                                            |
| Compact                 |  46k |     31k |   67k | The scroll is replaced by a summary a fraction of its size. That summary is written; the head is still read.                                      |
| After idling            |    0 |    230k |  460k | The cache entry is gone. Nothing is read; the entire window is written again at the write rate.                                                   |

_Prompt blocks in prefix cache order and how actions between turns affect cache reuse._

So, to reiterate: because the cache matches from left to right, simple actions can cause the entire window to be reread at full cost:

- Changing MCP config mid-session. Tool definitions sit at the very start of the prompt. Adding, removing, or updating an MCP tool busts the entire cache — the system prompt and all conversation history have to be rewritten.
- Changing the effort level or compacting. Instructions live inside the system prompt block. The tool definitions stay cached, but the instructions and every message that follows get rewritten.
- Switching models or speed modes. Fast mode alters the system prompt. Changing models drops the cache completely because cached key-value states cannot be shared between model architectures.

### Cache timeout

The upstream LLM providers generally cache calls originating from harnesses like Claude Code and Codex for 60 minutes. After that window of inactivity, the cache expires, and the entire conversation needs to be recomputed at full price.

As a reminder, re-reading a 300-thousand token window cold costs the same as **6 million cached reads** on average.

> **Tip: Track cache expiration**
>
> You don't need to guess when the cache goes cold. The JSON that Claude Code pipes into a custom [status line](https://skurekjakub.dev/blog/cc-statusline.md) carries a `context_window.prompt_cache` block with `expires_at`. Since `expires_at` updates on every API response, `ttl - (expires_at - now)` gives an exact countdown to expiration:
>
> ![Prompt cache expiration countdown via a custom status line](https://skurekjakub.dev/blog/tokenmaxxing/statusline.png)
>
> _Prompt cache expiration countdown via a custom statusline._
>
> (You can get this exact statusline at <https://github.com/skurekjakub/cc-statusline>)

## Compressing tool output

With tool execution making up nearly 45% of the [context](#context-composition) on average, filtering command output is another area worth having a look. This is the place where external tools are already widely adopted and extremely helpful.

Shell-level filters like `rtk` ([Rust Token Killer](https://github.com/rtk-ai/rtk)) sit between the harness and the shell, intercepting commands and trimming output before the model reads it. A couple of examples, taken from their documentation:

- **`git log -20`:** Strips commit hashes, diff stats, and metadata formatting, dropping the output from 122 KB to 5 KB (a 95% reduction, from \~30k tokens to 1.2k).
- **`npx vitest run`:** Collapses passing test suites into a one-line summary (`PASS (816) FAIL (0)`), cutting 235 bytes to 20 bytes (91% saved).
- **`ls -la`:** Removes repetitive file permissions and ownership columns, cutting output by 66%.

On individual commands, these filters cut output by 60% to 95%. But looking across 12,500+ commands (tool calls made by agents) in my RTK telemetry, the total session-wide saving was **29.9%**. Take note however that this only includes commands filtered by RTK, not every toolcall. From the 45% toolcall representation in the average context, this likely represents less than 5%.

![Token saving stats as output by the 'rtk gain' command.](https://skurekjakub.dev/blog/tokenmaxxing/rtk-stats.png)

The gap between 95% savings on individual commands and 28.9% overall comes down to **Read**. RTK never sees the requested file contents and has no way to compact it. And even if it did, doing that would be counterproductive — as a static program it doesn't have the authority or the knowledge to know what's safe to hide from the agent.

In harnesses that expose turn lifecycle hooks, you can influence **Read** behavior. For example, in Claude Code, you can prevent agents from accidentally dumping a large file into their context via the `PostToolUse` hook:

```bash
# Full implementation compacted for brevity.
# Have your agent reconstruct the code for its specific harness.

# Applies only to specified formats
file_path=$(printf '%s' "$input" | jq -r '.tool_input.file_path // ""')
case "$file_path" in
    *.ts|*.tsx|*.js|*.jsx|*.mjs|*.mts|*.cjs|*.cts) ;;
    *) exit 0 ;;
esac

# Skips files smaller that 6000 bytes ~= 1.5k tokens
min_bytes="${RTK_READ_HOOK_MIN_BYTES:-6000}"
original_bytes=$(wc -c < "$file_path")
[ "$original_bytes" -ge "$min_bytes" ] || exit 0

# If the model asked for specific lines, let it pass untouched
has_range=$(printf '%s' "$input" | jq -r '((.tool_input.offset // .tool_input.limit) != null)')
[ "$has_range" = "true" ] && exit 0

# Otherwise compact the signatures
compacted=$(rtk read "$file_path" --level aggressive 2>/dev/null)
```

I should note however, that most harnesses already come with a built-in protection for exactly this use case. In the event of a large file read, the output gets redirected to a scrath file in the agent's workspace and the harness notifies the agent. Use custom hooks if you want even stricter control over what enters the context.

Also, I should note that directly interfering with the model's workflow in ways that it didn't expect can be dangerous. If a filter strips out details the model actually needs — like a line number or a method body — the agent often assumes the command failed or gave partial output. It re-runs the command with custom flags or executes it directly (RTK allows a full bypass of its filters via `rtk proxy <command>`), burning more tokens than the filter saved in the first place.

## The session prefix

The second largest chunk of the average context, as identified in the introduction of this post. It contains, among others, the system prompt, schemas for default tools and subagents, MCP server definitions, and deferred custom skill and tool descriptions and definition.

A frequent recommendation online is disabling MCP servers or trimming rules to cut startup overhead. To see what difference that actually makes, I measured the first-turn footprint across a few environments:

**Interactive chart.** Before the first keystroke the harness reported 28k tokens in the window with this machine's settings and 25k with none; deferred rows are listed by name and not loaded.

| Part                                   | Category                        | With settings | No settings |
| :------------------------------------- | :------------------------------ | ------------: | ----------: |
| Ships with the harness                 | System prompt                   |          9.1k |        9.1k |
| Ships with the harness                 | Built-in tools + skills listing |           15k |         15k |
| From settings, plugins and memory      | Custom agents                   |           589 |           0 |
| From settings, plugins and memory      | Memory files                    |          1.3k |         922 |
| From settings, plugins and memory      | Messages                        |          1.3k |           8 |
| Referenced, schema kept out (deferred) | MCP tool schemas                |           21k |           0 |
| Referenced, schema kept out (deferred) | Built-in tool schemas           |           15k |         15k |

_Initial token footprint before typing, comparing default and bare environments._

With all user settings, hooks, memory files, and active plugins enabled, the harness reported around 28k tokens. In a completely bare environment with no custom configuration, that baseline dropped to \~25k initial tokens.

We can also edit the system defaults, such as the system prompt a bit, but cutting these settings usually costs more than it saves. Without project rules, the agent burns tokens reading source files just to discover build commands and conventions. Without skill descriptions, it cannot see available skills. With modified or trimmed skill descriptions there is a risk it uses them incorrectly, requiring retries. A fail case similar to what I described in the tool outputs section.

The chart also clearly shows that all custom MCP servers, skills, and plugins load fully deferred. The agent if aware of their existence only via their name and description, coming from each skill and tool's frontmatter. The heavy payload only ever hits the context when agents explicitly requests it via `ToolSearch` (<https://code.claude.com/docs/en/tools-reference>).

## Putting everything together

Once you understand that an agent session is append-only and that prompt caching is what makes repeated turns affordable, day-to-day session management writes itself. And as you grow comfortable with the basics, you will naturally incorporate additional tooling like RTK to further streamline your workflows.

Keeping an agent's context clean is not about micromanaging instructions or turning off tools. The startup prompt is a tiny fraction of total usage, and prompt caching takes care of the vast majority of warm turns.

What keeps sessions lean comes down to how you plan and work:

- **Compact before stepping away.** Run `/compact` before you leave your desk. If a session did go cold, use a cheaper model like Haiku to run the compaction first so you don't pay frontier model rates on a cold cache write.
- **Clear finished sessions.** Run `/clear` once a task is done. Keep pending work in plan files on disk instead of dragging an old session along.
- **Keep main session clean and focused.** Use shell filters like `rtk` on verbose commands, and consider hooks for repetitive file reads.
- **Offload exploration to subagents.** Let disposable workers read through files and run tests, write their findings to disk, and return only a short status line.

### Tokenmaxxing the session prefix

Now let's compare the impact pruning and optimizing your session prefix has against proper session hygiene.

To compare these changes side by side, I converted everything into **base input tokens** using Anthropic's pricing: cache reads cost 0.1× and 1-hour cache writes cost 2×. On that scale, the 1.18 billion tokens that went through the API when working on this website cost the same as 162 million uncached tokens:

| Cache state                 | Tokens sent | Rate | In base input tokens |
| :-------------------------- | ----------: | ---: | -------------------: |
| Read from cache             |      1,156M | 0.1× |                 116M |
| Written to the 1-hour cache |         23M |   2× |                  47M |
| Not cached                  |        0.2M |   1× |                 0.2M |
| **Total**                   |  **1,180M** |      |             **162M** |

Using the base input token rate, and knowing the number of API calls (each API call always contains the session prefix), we can derive the following across 6,340 turns:

| Change                                | What it did                                                                                                               | Share of total bill |
| :------------------------------------ | :------------------------------------------------------------------------------------------------------------------------ | :------------------ |
| Filter tool output (`rtk`)            | Cut \~29% of command output across the 43% tool slice                                                                     | \~10%               |
| Avoid idle cache timeouts             | Prevented 3.4M tokens of cold cache rewrites                                                                              | 4.0%                |
| Avoid editing rules or tools mid-task | Prevented 3.1M tokens of cache-busting rewrites                                                                           | 3.6%                |
| Compact sessions                      | Cost 1.2M rewrite tokens to summarise, but stops runaway context growth                                                   | 1.4%                |
| Strip all rules, skills, and memory   | Cut 13k tokens per call, leaving the agent without project context (initial cost difficult to determine, varies per task) | 6.3%                |
| Turn off all MCP servers              | Cut 1.1k tokens per call                                                                                                  | 0.5%                |
| Slim down `AGENTS.md`                 | Cut \~300 tokens per call                                                                                                 | 0.15%               |

Across the 6,340 API calls in this dataset, each prefix token (27k in this repo) was read thousands of times, and using:

```maths
base tokens per token sent   =   1,156.1M read     × 0.1
                               +   23.26M written  × 2
                               +    0.22M uncached × 1
                               ─────────────────────────
                                 1,179.6M sent

                             = 162.4M / 1,179.6M
                             = 0.138

base tokens per prefix token = 6,340 calls × 0.138
                             ≈ 870
```

Gives us a prefix token cost of roughly 870 base tokens. So the 27k initial context I used for comparison comes to around 24 million base tokens (`27k * 870`), or about **2%** of the overall traffic.

In other words, optimizing the prefix by **10%** optimizes the overall token traffic by **0.2%.** The usual tokenmaxxing advice — disabling MCP servers or cutting a few lines from your instructions — barely affects the overall stats. They are rounding errors compared to basic session hygiene and filtering terminal output. And while wiping out your project rules looks like a 6% win on paper, an agent with no instructions wastes far more tokens reading files to figure out your codebase.

## Extra credit: using subagents

This topic deserves a post on its own, but not touching on it here is impossible.

Contained to a single session, shell filters and careful session management only get you so far, especially on long-horizon workflows. If an agent needs to process and reason about data from dozens of large files, the context window fills up regardless of caching or command trimming.

The cleanest way to handle noisy exploration is delegating it to throwaway subagents. The following diagram visualizes a basic two-tier hierarchy:

![Architecture diagram showing control flow running vertically through status codes, while data flow runs horizontally through files on disk.](https://skurekjakub.dev/blog/agent-as-function/agent-as-function.drawio.svg)

_Control flow stays lean inside the context window; data flow moves horizontally through files on disk._

When you dispatch a task to a subagent, the harness starts an isolated context window. The subagent can grep through files, run test suites, and read verbose logs without polluting the parent session.

The key is having the subagent report back without injecting orchestrator context with unnecessary data:

1. Save subagent output to files. The worker writes its report, diffs, and findings to disk (such as `artifacts/task-report.md`).
2. Return only a short status line. The worker's final response in the main session should just be a status update and the path to the report file.

Let's look at an example from another repository. During a 24-hour autonomous run on another codebase (a multi-module data pipeline), an orchestrator coordinated 171 throwaway subagents across 1,800 turns. Those workers handled over 14,000 API calls and processed 1.54 billion tokens searching code, reading files, and running test suites. But the orchestrator itself only ran 16 direct file reads and 28 edits. Because intermediate logs and file dumps were discarded when each worker finished, the orchestrator stayed focused on coordinating agents and the overall task, rather than veering off-track due to a noisy context. A phenomenon also known as _context rot_.

**Interactive chart.** Across a 24-hour autonomous overhaul, the orchestrator context peaked at 838k tokens across 1,800 turns while up to 17 concurrent disposable workers processed 1542.4M cumulative tokens in isolated windows.

| Hour | Orchestrator Window | Orchestrator Calls | Active Workers | Worker Tokens (Hour) | Worker Tokens (Cum) |
| :--- | ------------------: | -----------------: | -------------: | -------------------: | ------------------: |
| 0h   |                520k |                174 |              6 |                32.5M |               32.5M |
| 1h   |                217k |                102 |             14 |                45.9M |               78.4M |
| 2h   |                299k |                 88 |             12 |                40.9M |              119.3M |
| 3h   |                391k |                 99 |             13 |                54.9M |              174.2M |
| 4h   |                444k |                 62 |             11 |                41.5M |              215.7M |
| 5h   |                513k |                 68 |             12 |                57.6M |              273.3M |
| 6h   |                575k |                 58 |             11 |                65.2M |              338.6M |
| 7h   |                655k |                 79 |             13 |                68.9M |              407.5M |
| 8h   |                747k |                 95 |             16 |                78.1M |              485.6M |
| 9h   |                838k |                122 |             17 |               120.1M |              605.6M |
| 10h  |                147k |                112 |             13 |                79.9M |              685.5M |
| 11h  |                206k |                 88 |             10 |                94.6M |              780.1M |
| 12h  |                238k |                 39 |             10 |               126.5M |              906.7M |
| 13h  |                243k |                 11 |              2 |                25.1M |              931.8M |
| 14h  |                419k |                151 |              4 |                48.2M |                980M |
| 15h  |                478k |                 64 |              2 |                57.9M |             1037.9M |
| 16h  |                196k |                 85 |              6 |                32.8M |             1070.6M |
| 17h  |                221k |                 34 |              5 |                51.7M |             1122.3M |
| 18h  |                310k |                 94 |              4 |                68.1M |             1190.4M |
| 19h  |                367k |                 58 |              6 |                  73M |             1263.5M |
| 20h  |                411k |                 50 |              5 |                57.8M |             1321.3M |
| 21h  |                442k |                 40 |              4 |                55.7M |             1376.9M |
| 22h  |                493k |                 60 |              6 |                85.2M |             1462.2M |
| 23h  |                531k |                 64 |              5 |                80.2M |             1542.4M |

_24-hour autonomous run: orchestrator context window vs worker subagent token throughput across 1,800 turns._

The basic idea here is that paying the fixed session prefix when initializing subagents is strictly lower than the cost from repeated main session compactions and subsequent rediscovery of established facts about the state of the given task.

---

Sources:

- [Anthropic Prompt Caching Documentation](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching)
- [Claude Code Documentation](https://code.claude.com/docs)
- [RTK (Rust Token Killer) Repository](https://github.com/rtk-ai/rtk)
- [RULER: What's the Real Context Size of Your Long-Context Language Models?](https://arxiv.org/abs/2404.06654)
- [NoLiMa: Evaluating Long-Context Information Association](https://arxiv.org/abs/2502.05167)
- [Chroma Research: Context Rot in Language Models](https://research.trychroma.com/context-rot)
