Twenty-eight model pairs, one comparable

measurementmodelsmethodology

Eight models cleared the volume floor, giving 28 pairs. Twenty-five never overlapped, two were underpowered, and the one testable pair came out null.

Eight models in my transcript corpus ran enough turns to be compared against each other. That gives 28 pairs. Twenty-five of them never ran in the same week, two overlapped but were too small, and the single pair that could be tested showed no separable difference. The corpus is 1,106 sessions, 7,175 human turns, 60 projects, 428 days, from node scripts/corpus.mjs over every Claude Code, Codex and Cursor transcript on this machine.

If you have adopted a new model recently and you have transcripts on disk, the tempting move is to compute rework rate per model and read off a winner. The numbers will come out. They will not mean what they look like.

The number that looks conclusive

Two rows from the model table, columns trimmed to fit:

claude-opus-5    525 turns 4w  rework 2%  23.4 tools/turn  40k/turn  led 2026-07-26→2026-08-15
claude-opus-4-8  448 turns 8w  rework 6%  25.2 tools/turn  94k/turn  led 2026-06-13→2026-08-05

Ten rework incidents in 525 turns against 25 in 448. Roughly a third of the rate, with the newer model ahead. That is the headline you would write.

Now look at when those turns happened. Weekly turn counts for the two models, from the timeline in corpus.mjs --json:

week beginningclaude-opus-4-8claude-opus-5
2026-06-0842—
2026-06-1573—
2026-06-22104—
2026-06-2975—
2026-07-0642—
2026-07-1318—
2026-07-20924
2026-07-27—63
2026-08-032220
2026-08-10—238

Two weeks contain turns from both. In the first, the old model ran 92 and the new one ran four. In the second, the new one ran 220 and the old one ran two. There is no week in which both did real work.

So this is not model A against model B. It is August against June and July: different repositories, different tasks, a different me. "Newer model, less rework" and "the work got easier over those weeks" are the same rows of data.

What the gate does

corpus.mts refuses the comparison rather than printing it with a footnote. Three thresholds, all in one place:

Turns outside the shared weeks are discarded for the pair, not down-weighted. If nothing survives, the pair gets a verdict instead of a number. Thirteen named models appear in the corpus; eight cleared the 40-turn floor. Of their 28 pairs, 25 had no shared week at all. Two had one shared week with too few turns on one side: 21 against 196, and 110 against 62.

The comment above the function states the problem directly:

A model release is confounded with time in the worst possible way: you adopt the new one and stop using the old one, so "new model, less rework" and "you got better at this over the same weeks" are the same data.

The one pair that cleared

claude-opus-4-8 against claude-fable-5. Three shared weeks — 2026-06-29, 2026-07-06, 2026-07-20 — with 209 turns on one side and 184 on the other. Rework 5.3% against 6.5%. The verdict:

· claude-opus-4-8 vs claude-fable-5: no separable difference across the 3 shared
  week(s) (z=-0.53) — 23 rework incidents between them is too few to call

Eleven incidents against twelve. A two-proportion z of -0.53, against a threshold of 2. That gap is what a coin flip produces at this sample size. Four hundred turns of genuinely concurrent work are not enough to rank two models on rework.

Every other pair gets this caveat, printed once with all 25 names (elided here):

- gpt-5.3-codex vs default; gpt-5.3-codex vs claude-opus-5; … ; composer-1 vs
  gpt-5.1-codex-max never ran at volume in the same week. Their rework rates come
  from different months, and rework was already falling for other reasons, so the
  difference between them cannot be attributed to the model.

One correction to my own output, since the point of this is not overclaiming: that caveat says rework "was already falling". In this corpus it was not falling. The trend line compares the first and last 15 weeks — 3,075 turns — and reports 5.1% against 5.3%. Flat. The confound runs in whichever direction the underlying drift happens to run, and here it barely runs at all. The conclusion holds either way. The wording is sloppier than the test.

The turns with no model at all

Half the corpus is not in the model table:

(no model ran)  3494 turns 43w  rework 6%  12.7 tools/turn  2k/turn  (not compared)

3,494 of 7,175 turns, 48.7%. Two unrelated causes land in that bucket. Some are requests aborted before any assistant record was written, which the code notes are about 30% rework by construction, because that is what aborting means. The rest are an instrumentation gap: Cursor is 61% of the turns here, and across 641 Cursor sessions a model name appears on 989 of 86,314 assistant records — 1.15%. Those turns ran 12.7 tool calls each and shipped real work.

They are excluded from every per-model rate. Folding them in would blame an abort on whichever model happened to be selected when escape was pressed. The count is not a measure of abandoned work. It is mostly a measure of what one harness records.

The same restraint shows up in the cost view. runs.mjs --cost reports 30.88B cache-read tokens against 126.5M output across 2,367 runs, and then declines to convert:

No currency is shown. The rate card is not part of this snapshot, and a dollar figure derived from an assumed price is an assumption rendered as a fact.

What you can measure instead

Your own transcripts describe your own behaviour well enough: where the turns went, which repository dominates the totals, which agent family burns the most context per run, how often you interrupt. None of that needs a control group.

Ranking models does, and normal adoption behaviour destroys the control group. If you want the comparison, build it on purpose: keep both models live for several weeks, split by something other than the calendar, and expect to need far more than a thousand turns before a few percentage points separate from noise.

Otherwise the honest output is the one above. Twenty-five pairs unattributable, two underpowered, one tested and null.