Twenty-eight model pairs, one comparable
Eight models cleared the volume floor, giving 28 pairs. Twenty-five never overlapped, two were underpowered, and the one testable pair came out null.
Eight models in my transcript corpus ran enough turns to be compared against each other. That gives 28 pairs. Twenty-five of them never ran in the same week, two overlapped but were too small, and the single pair that could be tested showed no separable difference. The corpus is 1,106 sessions, 7,175 human turns, 60 projects, 428 days, from node scripts/corpus.mjs over every Claude Code, Codex and Cursor transcript on this machine.
If you have adopted a new model recently and you have transcripts on disk, the tempting move is to compute rework rate per model and read off a winner. The numbers will come out. They will not mean what they look like.
The number that looks conclusive
Two rows from the model table, columns trimmed to fit:
claude-opus-5 525 turns 4w rework 2% 23.4 tools/turn 40k/turn led 2026-07-26→2026-08-15
claude-opus-4-8 448 turns 8w rework 6% 25.2 tools/turn 94k/turn led 2026-06-13→2026-08-05
Ten rework incidents in 525 turns against 25 in 448. Roughly a third of the rate, with the newer model ahead. That is the headline you would write.
Now look at when those turns happened. Weekly turn counts for the two models, from the timeline in corpus.mjs --json:
| week beginning | claude-opus-4-8 | claude-opus-5 |
|---|---|---|
| 2026-06-08 | 42 | — |
| 2026-06-15 | 73 | — |
| 2026-06-22 | 104 | — |
| 2026-06-29 | 75 | — |
| 2026-07-06 | 42 | — |
| 2026-07-13 | 18 | — |
| 2026-07-20 | 92 | 4 |
| 2026-07-27 | — | 63 |
| 2026-08-03 | 2 | 220 |
| 2026-08-10 | — | 238 |
Two weeks contain turns from both. In the first, the old model ran 92 and the new one ran four. In the second, the new one ran 220 and the old one ran two. There is no week in which both did real work.
So this is not model A against model B. It is August against June and July: different repositories, different tasks, a different me. "Newer model, less rework" and "the work got easier over those weeks" are the same rows of data.
What the gate does
corpus.mts refuses the comparison rather than printing it with a footnote. Three thresholds, all in one place:
MIN_MODEL_TURNS = 40— below this a model is listed but never entered into a pair.MIN_WEEK_SHARE = 20— a week counts as shared only if both models ran at least 20 turns in it.MIN_PAIR_ARM = 80— pooled across shared weeks, each side needs 80 turns before a rate is computed.
Turns outside the shared weeks are discarded for the pair, not down-weighted. If nothing survives, the pair gets a verdict instead of a number. Thirteen named models appear in the corpus; eight cleared the 40-turn floor. Of their 28 pairs, 25 had no shared week at all. Two had one shared week with too few turns on one side: 21 against 196, and 110 against 62.
The comment above the function states the problem directly:
A model release is confounded with time in the worst possible way: you adopt the new one and stop using the old one, so "new model, less rework" and "you got better at this over the same weeks" are the same data.
The one pair that cleared
claude-opus-4-8 against claude-fable-5. Three shared weeks — 2026-06-29, 2026-07-06, 2026-07-20 — with 209 turns on one side and 184 on the other. Rework 5.3% against 6.5%. The verdict:
· claude-opus-4-8 vs claude-fable-5: no separable difference across the 3 shared
week(s) (z=-0.53) — 23 rework incidents between them is too few to call
Eleven incidents against twelve. A two-proportion z of -0.53, against a threshold of 2. That gap is what a coin flip produces at this sample size. Four hundred turns of genuinely concurrent work are not enough to rank two models on rework.
Every other pair gets this caveat, printed once with all 25 names (elided here):
- gpt-5.3-codex vs default; gpt-5.3-codex vs claude-opus-5; … ; composer-1 vs
gpt-5.1-codex-max never ran at volume in the same week. Their rework rates come
from different months, and rework was already falling for other reasons, so the
difference between them cannot be attributed to the model.
One correction to my own output, since the point of this is not overclaiming: that caveat says rework "was already falling". In this corpus it was not falling. The trend line compares the first and last 15 weeks — 3,075 turns — and reports 5.1% against 5.3%. Flat. The confound runs in whichever direction the underlying drift happens to run, and here it barely runs at all. The conclusion holds either way. The wording is sloppier than the test.
The turns with no model at all
Half the corpus is not in the model table:
(no model ran) 3494 turns 43w rework 6% 12.7 tools/turn 2k/turn (not compared)
3,494 of 7,175 turns, 48.7%. Two unrelated causes land in that bucket. Some are requests aborted before any assistant record was written, which the code notes are about 30% rework by construction, because that is what aborting means. The rest are an instrumentation gap: Cursor is 61% of the turns here, and across 641 Cursor sessions a model name appears on 989 of 86,314 assistant records — 1.15%. Those turns ran 12.7 tool calls each and shipped real work.
They are excluded from every per-model rate. Folding them in would blame an abort on whichever model happened to be selected when escape was pressed. The count is not a measure of abandoned work. It is mostly a measure of what one harness records.
The same restraint shows up in the cost view. runs.mjs --cost reports 30.88B cache-read tokens against 126.5M output across 2,367 runs, and then declines to convert:
No currency is shown. The rate card is not part of this snapshot, and a dollar figure derived from an assumed price is an assumption rendered as a fact.
What you can measure instead
Your own transcripts describe your own behaviour well enough: where the turns went, which repository dominates the totals, which agent family burns the most context per run, how often you interrupt. None of that needs a control group.
Ranking models does, and normal adoption behaviour destroys the control group. If you want the comparison, build it on purpose: keep both models live for several weeks, split by something other than the calendar, and expect to need far more than a thousand turns before a few percentage points separate from noise.
Otherwise the honest output is the one above. Twenty-five pairs unattributable, two underpowered, one tested and null.