Cache-read is 96% of the token bill
Across 2,364 agent runs, cache-read was 96% of all tokens and output was 0.39%. Eight runs account for half the cache-read; one accounts for 22%.
Across 2,364 agent runs on one machine — Claude Code, Codex and Cursor — output tokens were 0.39% of all tokens counted. Cache-read was 96%, a ratio of 244 to 1. Output is the number every harness puts in front of you, and it is the part that does not matter. Eight of those 2,364 runs account for half of all the cache-read. One accounts for 22% on its own.
Here is the actual output of runs.mjs --cost against that corpus:
token composition
cache-read 30.88B 96% ████████████████████████████████████████████
cache-create 1.06B 3% ██
output 126.4M 0% █
The integers behind the bars are 30,878,931,963 cache-read, 1,057,155,064 cache-create and 126,440,624 output. They come from summing cache_read_input_tokens, cache_creation_input_tokens and output_tokens over every assistant record in every transcript. Nothing is modelled or estimated.
One honest note on the denominator. That three-way split leaves out plain uncached input, which was another 445M. Fold it in and cache-read is 95.0% rather than 96.3%. The shape does not change.
Cache-read is the conversation, replayed
Output is what the model wrote. Cache-read is everything it had to be handed in order to write it: system prompt, tool definitions, the file it read nine turns ago, every tool result since. All of it, re-sent on every turn.
That is why the totals look the way they do. A 40-turn session does not read its context once. It reads it forty times, and each read is larger than the last. Output grows with how much the agent says. Cache-read grows with turns multiplied by a context that is itself growing. By the time a long run finishes, the text the model produced is a rounding error against the text it consumed to produce it.
Eight runs, half the bill
Sorting all 2,364 runs by cache-read and taking a running total:
| top N runs | share of all cache-read |
|---|---|
| 1 | 22.1% |
| 5 | 43.0% |
| 8 | 51.7% |
| 50 | 82.6% |
| 100 | 87.1% |
The largest single run is one transcript file in one repository: 13,393 assistant records, 6,067 tool calls, 124 MB on disk, first record 19 June and last 26 July. It read 6.81 billion tokens of context.
The instructive part is what that run does not contain. Its largest single turn read 999,134 tokens — completely ordinary. There is no runaway turn to find, no anomaly to point at. It is 13,393 unremarkable turns at an average of 508k cache-read each, and the total only exists once you add them up.
Why a session view cannot show you this
Not because the number is hidden. Because it is unreadable one session at a time.
The cumulative figure is right there — session-viz prints cache read on its own session card, beside turns, tool calls and output tokens. It still tells you nothing on its own. Across Claude Code runs with token data, the median run read 992k tokens of context and the 90th percentile read 5.2M. A session showing 4M is either badly run or entirely normal, and nothing inside that session distinguishes the two.
The number becomes a fact only when you have several hundred comparable runs and can see where they sit in the distribution. That is a fleet-level view by construction. No session view, ours included, can give it to you.
The spread between families
Grouping the 1,217 subagent runs by what kind of agent they were:
| family | runs | cache-read/run | output/run | × the leanest |
|---|---|---|---|---|
| other | 769 | 2,133,466 | 15,378 | 2.1 |
| note-writer | 46 | 1,199,799 | 12,154 | 1.2 |
| scoped-editor | 129 | 1,044,042 | 8,206 | 1.0 |
| verifier | 273 | 999,265 | 8,126 | 1.0 |
Same harness, same machine. What differs is the prompt each family starts from and the tools it reaches for. One class of agent reads twice the context per run that a comparable one does, and that is a prompt you can go and edit.
Now the part that matters for honesty. Narrowing other to match the leanest family would save about 872M tokens: 41% of all subagent cache-read, but 2.8% of the corpus total, because subagents are only 6.8% of cache-read overall. The bulk sits in long interactive sessions, which this grouping does not split. Real effect, correctly measured, and smaller than the 2.1× headline makes it sound. The family boundaries come from a first-message heuristic, and the per-run figures above are rounded means, so treat the 872M as an order of magnitude.
Why there is no dollar figure
You will not find currency anywhere in this tool. It prints its own reason:
No currency is shown. The rate card is not part of this snapshot, and a dollar figure derived from an assumed price is an assumption rendered as a fact.
Turning 30.88B cache-read into money needs a per-token rate, and that rate depends on your plan, your model mix, your discounts, and which of those applied in which week across a 428-day span. Any single number would be wrong for most readers and unfalsifiable for all of them. It would also be the number everyone quotes.
Token counts are measured. Money is inferred. We print the measured thing and let you apply your own rate card, because you know it and we do not.
What is not in these numbers
- Cursor: 641 runs, zero cache-read recorded. Its stored token data is partial, so that zero is a measurement gap, not a lean agent. Every total above is a floor.
- Cloud sessions. Transcripts from claude.ai/code are not stored locally and are absent entirely.
- Per-harness rates are not comparable. Claude Code averaged 20.4M cache-read per run over 1,285 runs; Codex averaged 10.6M over 438. Different workloads, and we do not present that as a like-for-like result.
Every figure here came from two commands on one laptop: runs.mjs --cost for the composition and the family split, runs.mjs --json for the raw integers behind them. The corpus is 60 projects over 428 days, and it is mine. Run the same commands and see whether your ratio looks like ours. One machine is one machine.