Claude Code Token Waste: 9 Patterns Burning Your Budget (and the Fix) (2026)

Claude Code Token Waste: 9 Patterns Burning Your Budget (and the Fix) (2026)
Most Claude Code token waste is structural, not behavioral. The expensive part is not the code Claude writes. It is the context re-sent on every turn: a bloated CLAUDE.md, an uncapped chat, hook output stitched in per prompt, and a prompt cache that expired while you were reading. Anthropic's docs spell out every one of these mechanisms, and each has a fix that takes under a minute. This is the finite list of nine patterns, the mechanism and 30-second fix for each, plus what does not actually move the needle.
TL;DR
Token waste in Claude Code is mostly re-sent context, not generated code.
A bloated CLAUDE.md reloads in full every turn for the whole session.
The prompt cache expires after 5 minutes by default and silently re-bills.
MCP schemas, skills, and hooks add tokens you never asked for.
Nine patterns, nine fixes, each under 30 seconds. Audit tonight.
<your real name>, Founder at Vantaige · Published 2026-05-19 · 14 min read · Last reviewed 2026-05-19
Why is Claude Code burning so many tokens on my project?
Claude Code burns tokens because every turn re-sends your startup context: system prompt, CLAUDE.md, memory, skill descriptions, and any hook output. Anthropic's context-window simulation shows the system prompt alone is roughly 4,200 tokens and startup load can fill about 20% of a 200,000-token window before you type a word. The code Claude writes is rarely the expensive part.
The mental model that matters: a long session is not one request. It is hundreds of requests, each re-billing the context that did not change. A 4,800-token CLAUDE.md is not 4,800 tokens once; across a 120-turn session it is processed 120 times. The prompt cache exists to discount that repetition, which is why pattern 4 is the one most operators underestimate.
Below, each pattern lists what it is, why it costs tokens, the 30-second fix as a real command or path, and an illustrative before/after. The numbers are realistic examples to show the shape of the saving, not a claim that Vantaige ran a 430-hour study. The mechanisms come directly from the Anthropic prompt-caching docs and the Claude Code memory docs.
What are the 9 token-waste patterns in Claude Code?

The nine patterns are: a bloated CLAUDE.md, an uncapped conversation history, hook injection overhead, prompt-cache expiration, over-eager skills, MCP schema loading, global extended thinking, letting wrong answers finish, and SessionStart hook bloat. They are ordered roughly by how often they show up in real setups and how much each one repeats per turn.
# | Pattern | Where it lives | Illustrative before / after |
|---|---|---|---|
1 | Bloated CLAUDE.md | ./CLAUDE.md, ~/.claude/CLAUDE.md | 4,800 to 900 tokens/turn |
2 | Uncapped chat history | The conversation itself | 40k+ to 9k tokens/turn |
3 | Hook injection overhead | UserPromptSubmit hooks | 600 to 0 tokens/prompt |
4 | Prompt-cache expiration | 5-min default TTL | 10x re-bill to 0.1x read |
5 | Over-eager skills | Skill descriptions at startup | 3,200 to 400 tokens/turn |
6 | MCP schema loading | Connected MCP servers | 11,000 to ~120 tokens/turn |
7 | Global extended thinking | Thinking budget setting | 2,500 to 0 thinking tokens |
8 | Letting wrong answers run | The response stream | 3,000 wasted to interrupted |
9 | SessionStart hook bloat | SessionStart hooks | 1,400 to 150 tokens/session |
1. A bloated CLAUDE.md that reloads every single turn
CLAUDE.md is the persistent instruction file Claude Code reads at the start of every session. The cost is not the read. It is that the file sits in context all session and gets re-processed every turn. Anthropic's docs state plainly: "Files over 200 lines consume more context and may reduce adherence."
Why it costs tokens: CLAUDE.md is loaded in full regardless of length, and unlike auto memory there is no 200-line cap trimming it. A 4,800-token file of onboarding prose, changelog notes, and "remember to be nice" filler is re-sent on turn 1 and turn 120 identically. It also survives /compact by being re-read from disk, so compaction does not shrink it.
The 30-second fix: run /memory, open the project CLAUDE.md, and move anything that is not an always-true rule into a path-scoped rule. Put long procedures in a skill, and topic notes in .claude/rules/ with a paths: frontmatter block so they only load when Claude touches matching files.
# .claude/rules/api.md
---
paths:
- "src/api/**/*.ts"
---
# API rules (load only when API files are open)
- All endpoints validate input
- Use the standard error envelope
Illustrative before/after: a 4,800-token CLAUDE.md trimmed to a 900-token core (build commands, conventions, hard rules) saves roughly 3,900 tokens on every turn. Across a 100-turn session that is the same paragraph not being re-billed 100 times. See our companion piece on what to cut from context first for the prioritization order.
2. Conversation history re-read on every message
Every message you send re-sends the entire prior conversation. That is how the model stays coherent, but it means a 60-turn debugging session where turn 5 dumped a 12,000-token log is still carrying that log at turn 60. The conversation is the single largest growing cost in a long session.
Why it costs tokens: there is no automatic pruning of mid-conversation tool output until you compact. A noisy npm test dump, a 2,000-line file read once, an MCP response never needed again: all of it rides on every subsequent turn at full price.
The 30-second fix: two moves. Run /compact when a phase of work ends so the verbatim history collapses into a structured summary. And start a fresh session per task instead of one marathon chat. For the recurring pattern, edit or delete the prior message that dumped the noise rather than letting it ride. Compaction keeps your intent, key files, and fixes; it drops the raw tool spew.
Illustrative before/after: a session carrying 40,000+ tokens of stale tool output, compacted at a phase boundary, drops to roughly a 9,000-token summary. Every turn after that is billed against 9k, not 40k. Pairing this with subagents (which get a fresh, separate context) is covered in three subagent context patterns.
3. Hook injection overhead on every prompt
A UserPromptSubmit hook runs before Claude sees your message, and its stdout is added to the conversation as context. Anthropic's hooks docs are explicit: "any non-JSON text written to stdout is added as context." A hook that prints git status, the current time, or a 40-line "project state" banner is adding those tokens to every prompt you ever send.
Why it costs tokens: this is not one-time. It fires per prompt. A 600-token status banner injected on 80 prompts is 48,000 tokens on text you mostly do not need, and it breaks the message-level prompt cache because the messages block changed.
The 30-second fix: open .claude/settings.json, find the UserPromptSubmit hook, and either remove it or trim its stdout to one line. If the data is occasionally useful, gate it: only print when something actually changed, not on every prompt. Verbose context belongs in a skill you invoke on demand, not a per-prompt hook.
// .claude/settings.json: print nothing unless there is a real signal
"hooks": {
"UserPromptSubmit": [
{ "hooks": [ { "type": "command",
"command": "git status --porcelain | grep -q . && echo 'uncommitted changes present' || true" } ] }
]
}
Illustrative before/after: a 600-token-per-prompt banner reduced to zero on routine prompts saves 600 tokens every message and restores the message-level cache. Over an 80-prompt session that is roughly 48,000 tokens.
4. Prompt-cache expiration: the 5-minute miss that re-bills everything
This is the pattern operators underestimate most. Anthropic's prompt-caching docs state: "By default, the cache has a 5-minute lifetime. The cache is refreshed for no additional cost each time the cached content is used." If you read a doc, get coffee, and come back 7 minutes later, the cache expired. Your whole CLAUDE.md and system prompt get re-billed at full input price instead of the cached read price.
Why it costs tokens: cache read tokens are 0.1x the base input price, and a hit refreshes the timer for free. A miss after a pause means you pay 1x again on the entire cached prefix, then 1.25x to re-write the cache. The gap between "1x re-bill" and "0.1x read" on a large prefix is the silent multiplier on a bursty workday.
The 30-second fix: for API and SDK workloads where prompts recur less often than every 5 minutes, switch the cache to the 1-hour TTL by adding "ttl": "1h" to the cache_control block on the stable prefix. For interactive Claude Code use, the practical fix is to keep the cached prefix small (fix patterns 1, 5, and 6) so a miss costs less, and to work in focused bursts rather than scattered single prompts across an hour.
{
"type": "text",
"text": "Long stable system prefix",
"cache_control": { "type": "ephemeral", "ttl": "1h" }
}
Illustrative before/after: a 6,000-token cached prefix that misses on every coffee break re-bills at 1x repeatedly. Kept warm or moved to a 1-hour TTL, the same prefix reads at 0.1x. The 1-hour write costs 2x base once, then every read is 0.1x, which pays back fast on a bursty schedule.
5. Skills that auto-load "just in case"
Skill descriptions are listed in your startup context so Claude knows what it can invoke. The one-line descriptions are cheap individually, but a stack of 30 verbose skill descriptions adds up, and every one of them rides in context for the whole session whether or not you ever invoke it. Critically, the docs note this listing is "not re-injected after /compact", so only skills you actually used get preserved.
Why it costs tokens: bloated multi-sentence descriptions are the issue, not skills existing. Full skill content only loads on use, which is correct. But 30 skills with three-sentence descriptions is a few thousand tokens of menu sitting in every turn before compaction.
The 30-second fix: two moves. Tighten every skill description in its SKILL.md frontmatter to one tight sentence. And for skills with side effects or rare use, set disable-model-invocation: true so the description stays out of the startup index entirely until you invoke it with /name.
# skills/deploy/SKILL.md
---
name: deploy
description: Deploy current branch to production. Run only when asked.
disable-model-invocation: true
---
Illustrative before/after: 30 skills with verbose descriptions at roughly 3,200 tokens of startup menu, trimmed and with rare ones flagged disable-model-invocation: true, drops to around 400 tokens. That is per turn until the first compaction.
6. Connected MCP servers loading every tool schema
Connected MCP servers can inject the full JSON schema of every tool they expose into context. A handful of MCP servers with 40+ tools between them is one of the largest avoidable startup costs, because tool schemas are verbose and they sit in context for the whole session.
Why it costs tokens: Anthropic moved the default toward deferred tool loading for exactly this reason. The docs describe MCP tool names listed while "full schemas stay deferred and Claude loads specific ones on demand." If you forced full schema loading, every request carries thousands of tokens of tool definitions you mostly never call.
The 30-second fix: keep tool search on so schemas load on demand. Do not set ENABLE_TOOL_SEARCH=false (that forces everything to load upfront). Use ENABLE_TOOL_SEARCH=auto so schemas only load eagerly when they fit within 10% of the context window. And disconnect MCP servers you are not using this session rather than leaving every server wired.
# .claude/settings.json
"env": { "ENABLE_TOOL_SEARCH": "auto" }
Illustrative before/after: full schema loading for several MCP servers at roughly 11,000 startup tokens, switched to deferred loading, drops to the tool-name listing of around 120 tokens until a tool is actually needed. If MCP servers expose zero tools at all, that is a different bug, fixed in why your MCP server shows no tools.
7. Global extended thinking left on for trivial tasks
Extended thinking spends extra tokens on a reasoning pass before the answer. That is the right trade for a hard refactor or tricky bug. It is pure waste for "rename this variable." If thinking is on globally, you pay the reasoning tax on every trivial turn.
Why it costs tokens: thinking tokens are generated tokens billed at output rates, and they change the message cache, so a thinking pass on a simple prompt costs both reasoning tokens and a cache disruption. Across a day of small edits that is a steady leak with no quality benefit.
The 30-second fix: turn extended thinking off as the default and reach for it explicitly. In an interactive session, toggle it per task rather than leaving a global thinking budget on. Reserve it for prompts where you genuinely want a plan before the edit; leave it off for mechanical changes.
Illustrative before/after: a global thinking budget producing roughly 2,500 reasoning tokens on a one-line rename, switched to off-by-default, saves those 2,500 output-priced tokens per trivial turn and keeps the message cache intact. The model still answers; it just stops over-thinking a trivial ask.
8. Letting a wrong response finish instead of interrupting
When Claude starts down the wrong path, the most expensive thing you can do is wait politely for it to finish a 3,000-token answer, then correct it. You paid for the wrong output, the correction, and the re-answer. The cheapest fix in this entire list is a single keystroke.
Why it costs tokens: output tokens are the most expensive class. A wrong 3,000-token response is 3,000 output-priced tokens of pure waste, plus it sits in conversation history re-billing on every later turn until you compact, compounding pattern 2.
The 30-second fix: the moment the answer is clearly off, press the interrupt: Cmd+. on macOS, Ctrl+. on Windows and Linux. Then redirect with one tight sentence. Interrupting at token 200 instead of token 3,000 saves the other 2,800 output tokens and keeps the bad output out of your history.
Illustrative before/after: a wrong answer caught at token ~200 instead of running to ~3,000 saves roughly 2,800 output-priced tokens immediately, and avoids carrying a dead 3,000-token block forward on every subsequent turn.
9. SessionStart hook bloat: the "plugin loaded successfully" tax
A SessionStart hook prints to stdout, and that stdout is added to context at the start of the conversation. The docs confirm: "Any text your hook script prints to stdout is added as context for Claude." Five plugins each announcing "loaded successfully" with a banner is context you paid for and the model does not need.
Why it costs tokens: SessionStart output enters startup context and therefore rides every turn until the first compaction, exactly like CLAUDE.md. "Plugin X v2.1 loaded successfully" times five, with ASCII art, is a real number of tokens that buys zero capability.
The 30-second fix: open .claude/settings.json, find SessionStart hooks, and silence the cosmetic ones. A hook that only needs to load real context should print just that context; a hook that only does setup should print nothing on success and exit 0. Move "loaded" confirmations to stderr or drop them.
// after: do the setup, say nothing on success
"hooks": {
"SessionStart": [
{ "hooks": [ { "type": "command",
"command": "scripts/warm-cache.sh >/dev/null 2>&1 || true" } ] }
]
}
Illustrative before/after: five plugins printing ~1,400 tokens of startup banners, silenced, drops to around 150 tokens of genuinely useful loaded context, saved on every turn until compaction.
What did not move the needle?
Several popular "token saving" tactics produce little or no real reduction, and chasing them wastes time better spent on the nine patterns above. The honest answer is that the big wins are structural (re-sent context), not stylistic.
Writing terser prompts. Shaving your prompt from 60 words to 30 saves about 40 tokens once. It does nothing about the 6,000-token prefix re-sent every turn. The leak is the prefix, not your typing.
Asking the model to "be concise." It trims some output, but output is a fraction of a long session's cost versus re-sent context and tool history. Real output control comes from interrupting wrong answers (pattern 8), not politeness.
Switching to a smaller model for everything. This lowers per-token price, not token count. A bloated CLAUDE.md on a cheap model is still re-processed every turn. Fix the structure first.
Splitting CLAUDE.md into many
@imports. The docs are explicit that imported files "still load and enter the context window at launch." Imports help organization, not context size. Only path-scoped rules and skills defer the load.Clearing context constantly. Over-compacting destroys useful state and forces re-exploration, which re-reads files and spends more than it saved. Compact at phase boundaries, not reflexively.
How do I audit my own Claude Code token usage?

Run a five-minute audit in this order: open /memory to see exactly what is loading, line-count your CLAUDE.md, list your skills, check .claude/settings.json for UserPromptSubmit and SessionStart hooks, and list connected MCP servers. The biggest wins are almost always patterns 1, 2, and 6, because those carry the most tokens per turn.
See what loads. Run
/memoryin a session. It lists every CLAUDE.md, CLAUDE.local.md, and rules file actually loaded. Anything listed is being re-sent every turn.Measure CLAUDE.md. Run
wc -l CLAUDE.md ~/.claude/CLAUDE.md. Over 200 lines is the documented adherence-and-context threshold. Trim to a hard-rules core.Count skills. List your skills directory. For each, ask "does Claude need to know this exists every session." If not, set
disable-model-invocation: true.Inspect hooks. Open
.claude/settings.json. Read every UserPromptSubmit and SessionStart command. If it prints to stdout on success, it is taxing you.Audit MCP. Confirm
ENABLE_TOOL_SEARCHis unset orauto, neverfalse. Disconnect unused MCP servers. CLAUDE.md and skills grow back, so re-run this monthly.
Does the prompt cache really expire after 5 minutes?
Yes. Anthropic's documentation states the default cache lifetime is 5 minutes, refreshed for free on each use, with an optional 1-hour TTL available at higher write cost. A cache hit resets the timer at no charge; a miss re-bills the entire cached prefix at full input price plus the 1.25x cost to re-write the 5-minute cache.
This is why bursty work quietly costs more than focused work. Ten scattered prompts across an hour can each land after the cache expired, paying 1x on the full prefix every time. The same ten in one 8-minute block keep the cache warm and pay 0.1x reads after the first write. The fix is structural (small prefix) plus behavioral (work in bursts), or the explicit 1-hour TTL for API workloads.
Frequently asked questions
How much can fixing these patterns actually save?
It depends entirely on your setup, so treat these numbers as illustrative shapes, not a guarantee. The structural fixes (patterns 1, 2, 5, 6) compound because they cut tokens re-sent on every turn. A prefix dropping from roughly 20,000 startup tokens to 5,000 is not a one-time saving; it is 15,000 fewer tokens billed on every turn until compaction, then again after.
Will trimming CLAUDE.md make Claude follow my rules less?
The opposite, per Anthropic's docs: longer files "consume more context and reduce adherence." A focused 900-token CLAUDE.md of concrete, verifiable rules ("use 2-space indentation", "run npm test before commit") produces more consistent behavior than a 4,800-token file mixing rules with prose. Move detail to path-scoped rules and skills; keep CLAUDE.md to always-true facts.
Is the 1-hour cache always worth it over the 5-minute default?
No. The 1-hour cache write costs 2x base input price versus 1.25x for the 5-minute write. It pays back only when prompts recur less often than every 5 minutes but more than every hour, or when post-pause latency matters. For tight interactive loops the 5-minute default plus a small prefix is cheaper. Match the TTL to your actual request cadence.
Do skills cost tokens even if I never invoke them?
Their one-line descriptions do, because the description list is in startup context so Claude knows what it can invoke. Full skill content loads only on use. The fix is two-part: keep descriptions to one tight sentence, and set disable-model-invocation: true on rare or side-effecting skills so the description stays out of the index until you call it with /name.
Why does my token usage not drop after /compact?
Because CLAUDE.md, user instructions, and rules are re-read from disk and re-injected after /compact by design, so a bloated CLAUDE.md survives compaction unchanged. Compaction collapses the conversation, not your instruction files. To cut post-compaction cost you must trim CLAUDE.md and rules themselves, which is pattern 1.
Does interrupting a response actually save money?
Yes, and it is the cheapest fix here. Output tokens are the priciest class. Pressing Cmd+. (macOS) or Ctrl+. (Windows and Linux) at token ~200 instead of letting a wrong ~3,000-token answer finish saves roughly 2,800 output-priced tokens immediately, and stops that dead block from re-billing on every later turn until you compact.
Are MCP tool schemas loaded on every request?
By default Claude Code lists MCP tool names and defers full schemas, loading specific ones on demand. They are loaded eagerly only if you set ENABLE_TOOL_SEARCH=false or run an older default. Keep it unset or auto so schemas load only when a task needs them or when they fit within 10% of the context window.
What is the single highest-impact fix?
For most setups it is a tie between trimming CLAUDE.md (pattern 1) and capping or compacting the conversation (pattern 2), since both cut tokens re-sent every turn. MCP schema deferral (pattern 6) is the biggest single startup drop if you run several MCP servers.
Related from Vantaige
References
Anthropic, Prompt caching (5-minute default TTL, 1-hour TTL, cache write 1.25x, read 0.1x, cache invalidation hierarchy).
Claude Code docs, How Claude remembers your project (CLAUDE.md loaded full at session start, "Files over 200 lines consume more context and may reduce adherence", path-scoped rules, survives /compact).
Claude Code docs, Explore the context window (startup load breakdown, system prompt size, skill index not re-injected after /compact, what survives compaction).
Claude Code docs, Hooks reference (UserPromptSubmit stdout added as context, SessionStart stdout added as context, 10,000-character output cap).
Claude Code docs, Skills (skill descriptions in startup index, disable-model-invocation behavior).
Get the best new AI tools and guides, weekly
One short email a week. The tools worth trying, the guides worth reading, nothing else.
No spam. Unsubscribe anytime.
Aymen B
Contributing writer at Vantaige, covering the AI tools ecosystem.


