skurekjakub.dev

Interactive lab

Attention budget

When an LLM writes its next word, it has a fixed 100% attention budget to divide across your entire conversation history. This interactive playground shows how your prompts, context length, and chat history compete for that focus.


How attention works

Tokens & the context window: Language models do not process raw text. They split everything into small units called tokens (roughly 3–4 characters or ¾ of a word). Every part of your session—system instructions, past user messages, tool outputs, files read, and assistant replies—sits in one continuous list called the context window. On every turn, this entire history is sent to the model all over again.

The 100% attention budget: To generate the next word, the model looks back across all earlier tokens to determine what matters. It scores every token and converts those scores into percentages that must add up to exactly 100%. Think of it like a fixed pie: every token you add to the context window takes a slice away from everything else. Attention is strictly a zero-sum game.

The simulated scenario: an 80k coding session

To see attention economics in action, this sandbox simulates an active coding agent session currently sitting at 80,000 tokens. As an agent works through a multi-step programming task, its context window accumulates several distinct layers of history:

  • Token 0 (Session start): System prompt and project guidelines (such as CLAUDE.md, ~1,500 tokens). Models anchor their generation here by default.
  • Buried in the middle: Three architectural decisions and notes recorded earlier in the session:
    • Note A (~15% depth, ~12k tokens in): An API retries policy (“retries stay at 3 with a 250 ms backoff”).
    • Note B (~47% depth, ~37k tokens in): Root-cause diagnosis of an elusive authentication bug (“the middleware drops the session cookie when the login route answers with a 302 redirect”). This is the default target note selected for retrieval.
    • Note C (~76% depth, ~60k tokens in): A code style convention (“no default exports in lib. Named exports only”).
  • Recent tail (Right before your prompt): 1,400 tokens of fresh tool output—raw compiler logs, terminal dumps, or test traces emitted by the agent's very last action.
  • Background noise (~95% of the window): Routine command dumps, file contents, and past conversational turns that have no relevance to the auth bug.

You are now dispatching your next prompt to the agent. The sandbox calculates where every fraction of the model's 100% attention budget lands across all 80,000 tokens when predicting the very next word.

Guided experiments to try

Use the interactive sandbox below to step through these scenarios and observe how attention shifts across the context window:

  1. See the default position bias (The bare U): Select continue in the prompt panel, then toggle log scale on the attention distribution chart. With no relevant keywords in the prompt, attention follows position alone: a massive spike at the recent prompt, an anchor bump at token 0, and the middle completely neglected. Note B receives almost 0% attention.
  2. Specific wording is your best tool: Step through vague → paraphrase → exact wording. Watch Note B's attention jump from under 1% to ~20%. The note did not move; only your prompt's phrasing changed. Exact keywords trigger strong query-key vector alignment.
  3. Multi-topic prompts dilute attention: Choose three things at once. Asking about three different topics in a single message splits the prompt's attention vector, weakening the match for each individual note. Asking one focused question at a time is mathematically far more effective.
  4. Fiddle with context length (Context dilution): Expand What is in the context window at the top of the sandbox. With exact wording selected in the prompt panel, click through the 12k, 80k, 200k, and 400k quick buttons. Even though your prompt and Note B remain identical, Note B's share steadily drops from ~22% down to ~5% because hundreds of thousands of background filler tokens divide the same 100% pie.
  5. Simulate context rot (Look-alike distractors): Inside the context window disclosure, drag Look-alike notes elsewhere up to 4. Four outdated hypotheses about the bug now steal as much attention as the real answer. This demonstrates why coding sessions derail when stale debugging attempts accumulate in the window.
  6. Why re-pasting works so well: Still inside the context window panel, check Re-paste Note B right before the prompt. The newly pasted copy at the end of the context captures nearly all the attention, while the older copy in the middle is ignored. This explains why scratchpads and re-stating key facts keep long sessions accurate.
  7. Test your own wording: Select your own ↓ and write a custom prompt. Watch the relevance bars to see how well your phrasing hits the buried notes, then refine your vocabulary to see the attention spike.
What is in the context window80,000 tokens

Here is the complete conversation history sent to the model on this step, arranged chronologically from oldest (token 0) to newest. In real-world coding agents, most of the context is filled with command output and file contents, with a few crucial bug reports and architectural decisions buried in between.

System prompt + CLAUDE.md: tokens 0–1,500sysNote A · retries decision: tokens 12,000–12,300ANote B · the auth bug: tokens 37,600–37,900BNote C · exports rule: tokens 60,800–61,100CRecent tool output: tokens 78,583–79,983The prompt: tokens 79,983–80,000token 0token 80,000 · latest prompt
  1. 0–1,5001,500 tokSystem prompt & repository rules loaded at the start of the session.System prompt and CLAUDE.md. Comment policy: JSDoc on every function. Never count things in prose. Run npm run verify before claiming done. Shell output goes through rtk.
  2. in between76,183 tokBackground tool outputs, file reads, and past edits. Irrelevant to the current question.
  3. 12,000–12,300300 tokNote A · retries decisionDecision: retries stay at 3 with a 250 ms backoff. Do not raise the retry count again.
  4. 37,600–37,900300 tokNote B · the auth bug — the target note the prompt is trying to retrieveAuth bug: the middleware drops the session cookie when the login route answers with a 302 redirect.
  5. 60,800–61,100300 tokNote C · exports ruleRule: no default exports in lib. Named exports only.
  6. 78,583–79,9831,400 tokMost recent tool output and command results from the last action.
  7. 79,983–80,00017 tokThe latest prompt. The model generates the next token from the end of this message.fix the auth bug: the middleware drops the session cookie on the 302

The timeline bar above is drawn to scale: grey areas represent background noise (command outputs, logs, unrelated files), while coloured slices mark specific saved notes. The prompt presets below try to ask about Note B (the auth bug), which sits right in the middle of the history. Use the controls below to change the environment:

What it does: Changes the total length of the conversation history. The notes keep their relative positions. Every added token is background filler, but because all tokens share the same 100% budget, filler tokens dilute the attention available for important notes.
What to watch: Watch Note B's attention percentage shrink as you dial context from 12k up to 400k, even with an identical prompt.

What it does: Scatters outdated hypotheses or earlier debugging attempts throughout the session (each 85% similar to Note B). Because they resemble the target, they siphon attention away from the real answer.
What to watch: Note B's share drops as the red “look-alike” distractors take over. This simulates “context rot”—long sessions filled with failed attempts make it harder for the model to retrieve the correct note.

What it does: Selects how distance penalises older tokens.
• KERPLE (Logarithmic): Standard for modern models. Attention drops quickly over the first few hundred tokens, then levels off, allowing middle tokens to still be found with strong keyword matches.
• ALiBi (Linear): A strict straight-line drop. Under ALiBi, older tokens drop off so steeply that anything beyond a few thousand tokens gets virtually 0% attention regardless of prompt wording.

What it does: Pastes a fresh copy of Note B right above your prompt message. Because this copy sits at the very end of the history, it avoids the distance penalty and captures almost the entire attention budget, while the older copy in the middle is completely ignored.
Why it matters: This explains why pasting relevant code or adding a quick recap directly into your prompt works so reliably in long chats.

Regions: Rest of the context window, System prompt + CLAUDE.md, Note A · retries decision, Note B · the auth bug, Note C · exports rule, Look-alike notes, Recent tool output, Re-pasted note B, The prompt itself, Token 0 (the sink).

Latest prompt

This is the message you are sending right now. As the model starts generating the next token, its focus is guided by how well your prompt's vocabulary matches earlier notes. Try the presets below or type your own message:

• continue / irrelevant: Contains no keywords related to the notes. The model relies entirely on default position bias (the plain U-curve).
• vague: Only weakly hints at the problem (“fix the login thing”).
• paraphrase: Describes the issue accurately, but using different terms (“redirect seems to lose the session”).
• exact wording: Uses the exact terms found in Note B (“auth bug”, “middleware drops the session cookie”, “302”).
• three things at once: Asks about Notes A, B, and C in one message, forcing the prompt to split its attention vector across multiple topics.

System prompt + CLAUDE.md
estimated match cos 0.00 → relevance boost +0.0 logits
Note A · retries decision
estimated match cos 0.00 → relevance boost +0.0 logits
Note B · the auth bug
estimated match cos 0.85 → relevance boost +6.8 logits
Note C · exports rule
estimated match cos 0.00 → relevance boost +0.0 logits

Keyword relevance match: The bars show how strongly your prompt matches each note based on shared keywords and synonyms. A score of 0 means no relation (the note blends into background noise), while 1.0 is a direct match. Notice how exact terminology produces a much higher relevance boost than vague phrasing.

Attention budget breakdown (must sum to 100%): These tiles show the percentage of attention allocated to each section of your session. Attention is a zero-sum game—gaining focus in one area means losing it elsewhere. The final tile (Effective focus spread) measures how concentrated the model's attention is: 1 means laser-focused on a single token, while 80,000 means completely unfocused and diluted across everything.

Note B (Target Bug) · 42.2k tokens back
20.8%
300 tokens
Current prompt · 17 tokens
23.1%
Self-attention on prompt words
Recent tool output · last 1,400
29.5%
Earned by recency bias alone
Token 0 (Attention sink)
5.11%
Reserved sink; not reading text
Background context & filler
21.5%
Old logs, files, and chat history
Effective focus spread
1.5k
Out of 80k total tokens

Attention distribution across the context window

Share of attention per 200-token slice (oldest on the left, current prompt on the right). Sums to 100%.

sysABCtail0%18.3%36.6%020k40k60k80kposition in context window (tokens) → newest token generates from right edge
Hover over the chart to inspect a 200-token slice

How to read this chart & how attention shifts:
• Horizontal axis: Timeline of your session, from token 0 (start of session) to the prompt at the far right.
• Vertical bars: The share of the 100% budget given to each 200-token block, coloured by region.
• The Position Bias baseline (The U-curve): Without keyword relevance, attention defaults to token position alone. Models naturally allocate heavy focus to recent tokens (recency bias) and an anchor boost at Token 0 (the attention sink), while the long middle sags into near-zero attention. Toggle log scale to reveal this natural U-curve across the entire 80k-token session.
• The Content Relevance shift (Keyword & semantic match): When your prompt shares exact keywords with an earlier note, the query-key dot product spikes. This lifts Note B directly out of the middle sag, shifting up to ~20% of the entire probability budget onto the note.

Breakdown by region: Token counts and total attention share per region. Notice how 1,400 tokens of meaningless recent tool output can easily capture more attention than a 300-token critical bug report simply because the tool output is recent!

RegionTokensShare
Recent tool output1,40029.5%
The prompt itself1723.1%
Rest of the context window76,18321.1%
Note B · the auth bug30020.8%
Token 0 (the sink)15.11%
Note A · retries decision3000.18%
System prompt + CLAUDE.md1,4990.13%
Note C · exports rule3006.9e-2%

Under the hood: Raw scores (logits) before softmax

The solid black line is the baseline position score. Coloured dots show the boost each note gets from matching your prompt.

-10-8-6-4-20246810sysABCtailpromptsinklogit z = content + recency + head bias020k40k60k80kThe black line is baseline position bias; the shaded band is background noise (±1 logit)

How raw scores become attention percentages:
• Logits: Before computing percentages, the model calculates a raw score (called a logit) for every token.
• The black line (Position baseline): The baseline score granted purely by token position. It peaks at the recent right edge, bumps up at Token 0, and sags across the middle (the U-curve).
• The grey band (Noise floor): Unrelated background tokens naturally scatter within ±1 logit around this baseline.
• The coloured dots (Content boost): When your prompt matches a note, that note climbs upward from the baseline by its keyword relevance score. Because the model converts logits to percentages exponentially (via softmax), a lead of just 3 logits provides ~20× more attention, and a lead of 7 logits provides ~1,000× more attention!

Token-level score breakdown: Inspect the exact arithmetic for a sample token in each region.
• Distance: How many tokens ago this appeared.
• Content match: Relevance boost earned from prompt keywords.
• Recency penalty: Score lost due to distance.
• Head / sink bias: Positional bonus for the start of the session.
• Logit z: Total raw score (sum of content + position biases).
• Per-token share: Individual token's share of the 100% budget.
• Span share: Combined attention across all tokens in that whole section.

Worked example with current settings: A single token in Note B gains +6.66 from prompt keyword relevance but loses 7.79 to distance decay, ending with a raw score of −1.13 logits. Meanwhile, a token in recent tool output gains almost no relevance (+1.32) but only loses 3.00 to distance, ending with −1.68 logits. The difference is 0.54 logits, meaning each individual token of Note B is 1.7× more influential than a recent tool token. However, because the recent tool output is 1,400 tokens long while Note B is only 300 tokens, their combined section shares come out at 29.5% vs 20.8%.

TokenDistanceContent matchRecencyHead biasLogit zPer-token shareSpan share
Note B (Target), middle token42,249+6.66−7.79+0.00−1.134.1e-2%20.8%
Recent tool output, middle token716+1.32−3.00+0.00−1.682.4e-2%29.5%
The prompt, middle token8+2.95−0.14+0.00+2.812.14%23.1%
Token 0 (Attention sink)79,999−0.76−8.56+13.00+3.685.11%5.11%
Advanced mathematical parameters (Under the hood)

These parameters govern the underlying distance decay curves, attention sinks, and head biases. Adjust them to simulate different model architectures and attention behaviours. Click reset to restore defaults.

The vector size per attention head. Higher dimensions allow exact keyword matches to reach higher peak scores (scaling with √d).

How aggressively distance penalizes older tokens. At 0, recency bias is eliminated and older tokens are just as accessible as recent ones. Higher values make older context fade faster.

The token distance where steep exponential decay transitions into gentle logarithmic decay.

Active when using ALiBi prior. Head 8 is the shallowest slope in an 8-head model; steeper heads can only attend to the last few dozen tokens.

Extra score given to the very first token (Token 0). Trained models park unused attention probability here so it doesn't bleed into random filler tokens.

Positional boost given to tokens near the start of the session, where system instructions live. Setting this to 0 flattens the left side of the U-curve.

How far into the context window the system prompt boost extends before fading.

How strongly the generating token attends to other tokens within the current prompt message.

How much semantic synonyms (e.g. “redirect” vs “302”) count compared to exact keyword matches. Real embeddings give partial credit for related terms.

How closely distractor notes resemble Note B. High values (0.85–0.95) represent confusing near-duplicates that split the attention budget.

The mathematics & literature sources

a_i = exp(z_i) / Σ_j exp(z_j)          z_i = q·k_i / √d  +  b(i)

Scaled dot-product attention (Vaswani et al. 2017, arXiv:1706.03762). The softmax turns raw token scores z_i into probabilities that sum to 1. Because the denominator sums across all tokens in the context window, every added token dilutes the available probability budget.

q·k / √d  =  √d · cos θ        unrelated pair ~ N(0, 1)      exact match = √d = 8 (d = 64)

Content scoring & vector scaling: When query and key vectors align, their dot product scales to √d (8 at d = 64, ~11.3 at d = 128). Unrelated token pairs produce random Gaussian noise centered at 0 with unit variance.

b_rec(i) = − r₁ · log(1 + r₂ · |N − i|)          (KERPLE, logarithmic decay)
b_rec(i) = − m · (N − i),   m = 2^(−8h/n)         (ALiBi, linear decay)

Distance penalties: Chi et al. 2022 (KERPLE, arXiv:2205.09921) and Press et al. 2021 (ALiBi, arXiv:2108.12409). KERPLE's logarithmic decay closely mirrors the empirical distance decay observed in trained models using Rotary Position Embeddings (RoPE).

b_sink(0) = + σ

Attention sinks: Xiao et al. 2023 (StreamingLLM, arXiv:2309.17453). LLMs allocate substantial attention to the initial tokens (Token 0) as an “attention sink” to safely absorb unneeded probability mass when no specific past token is relevant.

b_head(i) = + π · exp(−i / ρ)

Lost in the Middle / Primacy bias: Hsieh et al. 2024 (Found in the Middle, arXiv:2406.16008). Attention naturally forms a U-shaped curve across long contexts, prioritizing initial instructions and recent messages while sagging across the middle.

cos θ_note ≈ √( Σ_hit w / Σ_all w ),   w = 1 / (notes containing the word)

Keyword overlap & semantic matching: Modarressi et al. 2025 (NoLiMa, arXiv:2502.05167). Demonstrates how exact keyword overlap drives effective context retrieval, while purely semantic paraphrases suffer significant retrieval penalties.

effective tokens = exp( − Σ a_i · ln a_i )

Effective focus spread (Perplexity of attention): Standard information-theoretic measure calculated as the exponential of entropy. Quantifies the effective number of tokens across which attention is actively distributed.


Companion interactive model for Tokenmaxxing. Mathematical formulations and priors derived from published literature (Vaswani et al. 2017, Chi et al. 2022, Press et al. 2021, Xiao et al. 2023, Modarressi et al. 2025).