Cut Your AI Agent Token Bill 5x: The Claude Code Efficiency Patterns (2026)

Cut Your AI Agent Token Bill 5x: The Claude Code Efficiency Patterns (2026)
You can cut Claude Code token cost by a large multiple without changing what you build. The spend is mostly re-sent context and unnecessary reasoning, not the code itself. On one identical high-complexity task, independent testing reported Claude Code finishing in 33,000 tokens against a competing agent's 188,000, roughly a 5.5x gap. The same efficiency math applies inside Claude Code too. This guide is the positive companion to our token-waste audit: six patterns that materially lower the bill, with the exact mechanism and a real config for each.
TL;DR
Token cost is mostly re-sent context, not generated code.
Prompt caching reads at 0.1x base input price.
Route trivial work to cheaper models, reserve Opus.
Scoped tools and pruned context cut per-turn spend.
Six patterns stack to a roughly 5x reduction.
<your real name>, Founder at Vantaige · Published 2026-05-19 · 13 min read · Last reviewed 2026-05-19
Why does Claude Code use so many fewer tokens than other agents?
Claude Code uses fewer tokens because it sends a tighter context per turn and defers anything it does not need yet. On one reported identical high-complexity task, Claude Code completed in about 33,000 tokens while a competing agent used about 188,000 for the same result, a roughly 5.5x gap noted across several 2026 comparisons. The model is not cheaper per token; it processes less context.
That gap is not magic. It is the sum of design choices you can also tune yourself: a cached stable prefix, deferred tool schemas, on-demand skills, and verbose work pushed into separate contexts. Anthropic's own cost guidance frames the principle plainly: "Token costs scale with context size: the more context Claude processes, the more tokens you use." Every pattern below shrinks what gets processed per turn.
The mental model that matters: a long session is hundreds of requests, each re-billing the context that did not change. Cut the re-sent bytes once and you cut them on every turn for the rest of the session. This article focuses on the techniques that reduce spend. For the inverse list of what silently inflates it, see the 9 token-waste patterns and their fixes.
What are the Claude Code efficiency patterns that cut token cost?
The six patterns are: prompt caching, context pruning, model routing, scoped tools, batching reads into subagents, and stopping wrong answers early. Each attacks a different cost driver. Caching discounts the repeated prefix, pruning shrinks the growing conversation, routing lowers the rate on trivial turns, scoped tools strip startup bloat, batching keeps verbose output out of your window, and stop-early kills wasted output tokens.
Technique | Mechanism | Typical saving |
|---|---|---|
Prompt caching | Stable prefix billed at 0.1x on cache read instead of 1x | Up to 90% off the cached prefix per turn |
Context pruning | /compact and /clear collapse stale tool output | 40k+ to ~9k tokens carried per turn |
Model routing | Haiku or Sonnet for simple turns, Opus reserved | Roughly 5x to 12x lower rate per simple turn |
Scoped tools | Deferred MCP schemas, CLI over MCP, on-demand skills | ~11k startup tokens to ~120 until needed |
Batching to subagents | Verbose reads stay in subagent context, summary returns | 6,100 read tokens return as ~420 |
Stop-early | Interrupt a wrong answer before it finishes | ~2,800 output tokens saved per catch |
Numbers above and below are illustrative shapes drawn from Anthropic's documented mechanics and pricing, not a claim that Vantaige ran a private study. Your actual saving depends on your CLAUDE.md size, MCP servers, and session length. The point is the direction and the multiplier, which compound when you stack all six.
How does prompt caching cut Claude Code token cost?

Prompt caching cuts cost by billing your stable prefix at 0.1x the base input price on a cache hit instead of 1x every turn. Anthropic's prompt-caching docs state the cache "has a default 5-minute lifetime" and "is refreshed at no additional cost each time the cached content is used." The system prompt, CLAUDE.md, and tool definitions are the prefix that repeats, so caching them is the single biggest structural saving.
The pricing is explicit. A 5-minute cache write costs 1.25x base input, a 1-hour write costs 2x, and every read or refresh costs 0.1x. So the first turn pays 1.25x to write a large stable prefix, then every subsequent turn within the window pays 0.1x to read it instead of 1x. On a 6,000-token prefix across 100 turns, that is the difference between paying full price 100 times and paying it once.
Claude Code applies caching automatically to repeated content like the system prompt. The lever you control is keeping the cache warm and the prefix small. For API and SDK workloads where prompts recur less often than every 5 minutes, set the 1-hour TTL on the stable prefix:
{
"system": [
{
"type": "text",
"text": "Long stable system prefix that rarely changes",
"cache_control": { "type": "ephemeral", "ttl": "1h" }
}
]
}
One gotcha from the docs: the minimum cacheable prompt length is 4,096 tokens for Claude Opus 4.7, 4.6, 4.5 and Haiku 4.5, and 1,024 tokens for Sonnet 4.6 and 4.5. Shorter prefixes are not cached and no error is returned. Illustrative shape: a 6,000-token prefix that reads at 0.1x instead of re-billing at 1x on 100 turns turns roughly 600,000 prefix tokens into about 60,000 plus one write. The cache-expiry trap that undoes this is detailed in what to cut from context first.
How do I prune context to lower per-turn token cost?
Prune context by collapsing stale tool output with /compact at phase boundaries and starting fresh with /clear between unrelated tasks. Anthropic's cost docs are direct: "Stale context wastes tokens on every subsequent message." Every message re-sends the full prior conversation, so a 12,000-token log read at turn 5 is still billed at turn 60 until you compact it away.
Compaction replaces the verbatim conversation with a structured summary. Per Anthropic's context-window docs, the summary keeps "your requests and intent, key technical concepts, files examined or modified with important code snippets, errors and how they were fixed, pending tasks, and current work" while dropping raw tool spew. You keep the meaning, not the megabytes. Steer it to protect what matters:
# In CLAUDE.md, a compact-instructions block
# Compact instructions
When you compact, focus on test output, file paths, and code changes.
Drop verbose logs and intermediate reasoning.
Two habits do most of the work. Run /compact when a phase of work ends, not reflexively, because over-compacting forces re-exploration that re-reads files and spends more than it saved. And use /clear when switching to unrelated work so a fresh session starts at a clean prefix. Use /rename before clearing and /resume to return, as the docs recommend.
Illustrative shape: a session carrying 40,000+ tokens of stale tool output, compacted at a phase boundary, drops to roughly a 9,000-token summary. Every later turn is billed against 9k, not 40k. Pairing this with subagents that get their own context is covered in three subagent context patterns.
How does model routing reduce my Claude Code bill?
Model routing reduces the bill by matching model price to task difficulty instead of paying Opus rates for trivial edits. Anthropic's cost guidance is explicit: "Sonnet handles most coding tasks well and costs less than Opus. Reserve Opus for complex architectural decisions or multi-step reasoning." Routing does not cut token count; it cuts the rate per token on the turns that do not need the strongest model.
Three controls do this. Use /model to switch mid-session, set a default in /config, and for subagents that do mechanical work, pin the cheapest tier in the subagent frontmatter. A renaming or boilerplate task on Haiku costs a fraction of the same task on Opus, with no quality loss because the task was never hard.
# A mechanical subagent: pin the cheapest model
---
name: log-summarizer
description: Grep and summarize log files. Returns only the failing lines.
model: haiku
---
Extended thinking is part of routing. The docs note thinking "is enabled by default" and "thinking tokens are billed as output tokens, and the default budget can be tens of thousands of tokens per request." For simple work, lower it with /effort, disable it in /config, or cap it with MAX_THINKING_TOKENS=8000. Reserve deep reasoning for prompts that genuinely need a plan.
Illustrative shape: routing the roughly 70% of turns that are mechanical to a model and effort level several times cheaper per token, while keeping Opus for the hard 30%, lowers blended cost without touching output quality on the work that mattered. The full routing matrix across models is in the instant-vs-Opus routing matrix.
How do scoped tools and on-demand skills cut startup token cost?
Scoped tools cut startup cost by keeping verbose tool schemas and skill bodies out of context until a task actually needs them. Anthropic's context-window docs confirm "MCP tool names listed so Claude knows what is available. By default, full schemas stay deferred and Claude loads specific ones on demand." A handful of MCP servers with 40+ tools is one of the largest avoidable startup costs if you force full schema loading.
Three moves. Keep tool search on so schemas defer; do not set ENABLE_TOOL_SEARCH=false. Prefer CLI tools where they exist, since the docs note tools like gh, aws, and gcloud "are still more context-efficient than MCP servers because they don't add any per-tool listing." And disconnect MCP servers you are not using this session with /mcp.
# .claude/settings.json: let schemas load on demand
{
"env": { "ENABLE_TOOL_SEARCH": "auto" }
}
Skills follow the same principle. Skill descriptions sit in startup context so Claude knows what it can invoke, but full bodies load only on use. For skills with side effects or rare use, set disable-model-invocation: true so the description stays out of the startup index entirely until you call it with /name:
# skills/deploy/SKILL.md
---
name: deploy
description: Deploy current branch to production. Run only when asked.
disable-model-invocation: true
---
Illustrative shape: full schema loading for several MCP servers at roughly 11,000 startup tokens, switched to deferred loading, drops to the tool-name listing of around 120 tokens until a tool is needed. If an MCP server shows zero tools instead, that is a separate bug fixed in why your MCP server shows no tools.
How does batching reads into subagents save tokens?
Batching saves tokens by running verbose operations in a subagent's separate context window so only a short summary returns to your main conversation. Anthropic's cost docs put it directly: "Delegate these to subagents so the verbose output stays in the subagent's context while only a summary returns to your main conversation." The big file reads never touch the context you pay for on every later turn.
The arithmetic is in the context-window simulation. A research subagent reads 6,100 tokens of files across several reads, all in its own window, and returns a 420-token result plus a small metadata trailer. Your main context grows by 420, not 6,100, and that 420 is what gets re-billed each turn afterward, not the full read. The subagent also starts without your conversation history, so it carries no accumulated baggage.
Use it for the operations the docs name: "Running tests, fetching documentation, or processing log files." A hook can do the same upstream by preprocessing before Claude sees the data. The docs give the canonical case: "Instead of Claude reading a 10,000-line log file to find errors, a hook can grep for ERROR and return only matching lines, reducing context from tens of thousands of tokens to hundreds."
# A PreToolUse hook that filters test output to failures only
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{ "type": "command", "command": "~/.claude/hooks/filter-test-output.sh" }
]
}
]
}
}
Illustrative shape: a research task that would add roughly 6,100 tokens of file reads to your main window returns as a 420-token summary, and a noisy test run filtered by a hook drops from tens of thousands of tokens to hundreds. Both keep the expensive bytes off every subsequent turn.
Why is stopping a wrong answer early the cheapest token saving?
Stopping a wrong answer early is the cheapest saving because output tokens are the priciest class and a single keystroke kills the rest of a bad response. The moment Claude heads down the wrong path, the expensive move is waiting politely for it to finish a 3,000-token answer, then correcting it. You paid for the wrong output, the correction, and the re-answer.
Anthropic's cost docs name the habit: "Course-correct early: If Claude starts heading the wrong direction, press Escape to stop immediately. Use /rewind or double-tap Escape to restore conversation and code to a previous checkpoint." Plan mode prevents the wrong path in the first place: "Claude explores the codebase and proposes an approach for your approval, preventing expensive re-work when the initial direction is wrong."
Two adjacent habits from the same docs compound this. "Give verification targets" so Claude catches its own mistakes before you request fixes, and "test incrementally" so issues surface "when they're cheap to fix." Each one stops a wrong direction from generating tokens you will pay for and then carry forward in history until you compact.
This pattern compounds with context pruning. A wrong 3,000-token answer is not a one-time charge. It sits in conversation history and is re-billed on every subsequent turn until you compact, exactly like a stale log read. Interrupting at token 200 saves the 2,800 output tokens now and removes the per-turn tax later, so a single keystroke pays back twice.
Illustrative shape: a wrong answer caught at roughly token 200 instead of running to 3,000 saves about 2,800 output-priced tokens immediately and keeps a dead 3,000-token block out of every later turn. The interrupt is Cmd+. on macOS and Ctrl+. on Windows and Linux, or Escape to stop and rewind.
How much can these patterns actually save together?

Stacked, these patterns target the same order of magnitude as the reported 5.5x cross-agent gap, because they attack different cost drivers that multiply rather than add. Caching discounts the prefix to 0.1x, scoped tools and pruning shrink what gets sent, routing lowers the rate on most turns, and stop-early plus subagents cut wasted output. Treat the multiplier as the lesson, not a guarantee for your exact setup.
Walk one realistic session. A bloated startup of roughly 20,000 tokens trimmed to 5,000 via scoped tools and a tight CLAUDE.md is 15,000 fewer tokens every turn. Caching reads that 5,000-token prefix at 0.1x instead of 1x. Routing the mechanical 70% of turns to a cheaper model multiplies the saving on rate. Subagents keep 6,000-token reads out of the window. None of these conflicts; each lands on a different part of the bill.
The honest framing: the big wins are structural, not stylistic. Writing terser prompts saves about 40 tokens once and does nothing about the prefix re-sent every turn. The real saving is in caching, context size, model price, and killing waste before it generates, which is exactly the set above. Anthropic's own summary of the principle is that costs "scale with context size," so the entire strategy is keeping context small and cheap to re-read.
Order of operations matters when you apply these. Start with scoped tools and a trimmed CLAUDE.md, because that shrinks the prefix every other technique then operates on. Caching a 5,000-token prefix is worth far more than caching a 20,000-token one, and a small prefix also clears the 4,096-token minimum cleanly on Opus-class models. Add context pruning next so the conversation stops growing unbounded, then layer routing and subagents on top. Stop-early is the one habit that pays from the first session with no setup at all.
What did not move the needle?
Several popular tactics produce little real reduction, and chasing them costs time better spent on the six patterns. The pattern is the same as in the waste audit: the big wins are structural, not cosmetic.
Writing shorter prompts. Trimming a prompt from 60 to 30 words saves about 40 tokens once. It does nothing about the cached prefix or the conversation re-sent every turn.
Asking the model to be concise. It trims some output, but output is a fraction of a long session's cost versus re-sent context. Real output control is interrupting wrong answers, not politeness.
Switching to a cheap model for everything. This lowers rate, not token count, and a weak model on a hard task often loops and costs more. Route by difficulty instead of downgrading globally.
Splitting CLAUDE.md into many imports. Imported files still load into context at launch. Imports help organization, not size. Only path-scoped rules and on-demand skills defer the load.
Clearing context constantly. Over-compacting destroys useful state and forces re-exploration that re-reads files. Compact at phase boundaries, not on reflex.
Frequently asked questions
Is the 5.5x token gap between Claude Code and other agents real?
It is a reported figure, not a Vantaige measurement. Multiple 2026 comparisons describe Claude Code completing an identical high-complexity task in roughly 33,000 tokens while a competing agent used about 188,000, a gap near 5.5x. The number varies by task and tool version. The reliable takeaway is the cause: a tighter per-turn context and deferred loading, both of which you can tune yourself with the patterns here.
Does prompt caching really cut cost by 90 percent on the prefix?
On a cache hit, yes for the cached portion. Anthropic's docs price cache reads at 0.1x base input, so a stable prefix read from cache costs one tenth of re-billing it at 1x. The first turn pays 1.25x to write the 5-minute cache or 2x for the 1-hour TTL. Net, after the write, every in-window turn reads the prefix at 0.1x, which is the structural saving caching exists to deliver.
When should I use the 1-hour cache TTL instead of the 5-minute default?
Use the 1-hour TTL when prompts recur less often than every 5 minutes but more than every hour, or when post-pause latency matters. The 1-hour write costs 2x base input versus 1.25x for the 5-minute write, so it only pays back on a bursty cadence. For tight interactive loops the 5-minute default with a small prefix is cheaper. Match the TTL to your real request rhythm, not a guess.
Which model should I default to in Claude Code to save money?
Default to Sonnet for general coding and reserve Opus for hard architecture or multi-step reasoning, per Anthropic's cost guidance. For mechanical subagent work, pin model: haiku in the subagent frontmatter. Switch mid-session with /model when a task genuinely needs more capability. Routing by difficulty lowers the blended rate without hurting quality on the work that actually needed the stronger model.
Do subagents actually reduce my token bill or just move tokens around?
They reduce your main context's recurring cost. A subagent's large file reads live in its own window and only a short summary returns. In Anthropic's simulation, 6,100 tokens of reads come back as a 420-token result. Since your main context is re-billed every turn, keeping 6,100 tokens out of it and carrying 420 instead compounds across the rest of the session, which is the real saving.
What is the single highest-impact efficiency change?
For most setups it is shrinking the cached prefix, because caching plus a small prefix cuts the largest repeating cost on every turn. Trim CLAUDE.md to hard rules, defer MCP schemas, and move detail into on-demand skills, then let caching read the small prefix at 0.1x. Context pruning is a close second since it caps the other growing cost, the conversation itself.
Does extended thinking waste tokens on simple tasks?
It can. Anthropic's docs note thinking tokens "are billed as output tokens" and the default budget "can be tens of thousands of tokens per request." On a one-line rename that is pure waste. Lower it with /effort, disable it in /config, or cap it with MAX_THINKING_TOKENS=8000 for simple work, and keep full thinking for prompts that genuinely need a plan before the edit.
How do I see my actual Claude Code token usage?
Run /usage for session token statistics and an estimated cost, and /context for a live breakdown by category with optimization suggestions. Run /memory to see which CLAUDE.md and auto memory files loaded at startup. The console Usage page is authoritative for billing. Measure first, then apply the patterns to the categories that are actually large in your setup rather than guessing.
Related from Vantaige
References
Anthropic, Prompt caching (5-minute default TTL, 1-hour TTL, cache write 1.25x and 2x, read 0.1x, refreshed free on use, minimum cacheable length by model).
Claude Code docs, Manage costs effectively (context scales token cost, model routing, /clear and /compact, MCP deferral, CLI over MCP, hooks preprocessing, extended thinking, subagent delegation, plan mode, course-correct early).
Claude Code docs, Explore the context window (startup load breakdown, deferred MCP schemas, subagent separate context returning a summary, what survives compaction).
Claude Code docs, Subagents (isolate high-volume operations in a separate context window, choose a model per subagent).
Claude Code docs, Skills (skill descriptions in the startup index, full bodies load on use, disable-model-invocation behavior).
Get the best new AI tools and guides, weekly
One short email a week. The tools worth trying, the guides worth reading, nothing else.
No spam. Unsubscribe anytime.
Aymen B
Contributing writer at Vantaige, covering the AI tools ecosystem.


