August 25, 2026
Opus 4.8 and Opus 5 tied on strict functional score across 25 matched Stet tasks, but their paired artifacts differed: Opus 4.8 had the lower footprint on 20 tasks, while Opus 5 ran more shell commands on 18 and more tests on 15.
Claude Opus 4.8 · Claude Opus 5
July 20, 2026
Six interventions meant to save tokens gave different answers depending on how you count. None cut the ten-task batch total twice. Caveman and the cheaper Terra model looked mildly cheaper per task, but within run-to-run noise.
GPT-5.6 Sol · GPT-5.6 Terra
July 8, 2026
On 24 real tasks, Sonnet scaled effort into more checking while Opus stayed flatter through high. The graders leaned Sonnet on clarity and Opus on diff minimality. Here is when I would use each.
Sonnet 5 · Opus 4.8
June 19, 2026
GLM 5.2 vs Composer 2.5 and the premium field on 50 real merged PRs from graphql-go-tools (Go) and sqlparser-rs (Rust). GLM lands last on craft and equivalence in both repos, costs about twice Composer, and writes more code than the human while missing the change. A routing guide for the new cheap model.
GLM 5.2 · Composer 2.5 · Opus 4.8 · GPT-5.5
June 18, 2026
Composer 2.5 vs Opus 4.8, GPT-5.5 and Opus 4.7 on 50 real merged PRs from sqlparser-rs (Rust) and graphql-go-tools (Go). Composer is 6.5-7x cheaper and ties the test gate, but finishes last on craft in both repos. A routing guide for where it's safe.
Composer 2.5 · Opus 4.8 · GPT-5.5 · Opus 4.7
June 14, 2026
Fable 5 vs Opus 4.8, GPT-5.5 and two more on 30 real GraphQL Go Tools and SQLParser Rust tasks. The useful read is taskwise W/D/L: which model won which metric, on which task, and why.
Fable 5 · Opus 4.8 · Opus 4.7 · GPT-5.5 · Composer 2.5
June 2, 2026
I graded four frontier coding models on 50 real merged PRs in Go and Rust - not just whether tests pass, but craft, equivalence, and cost. Opus 4.8 led on craft in both.
Claude Opus 4.8 · Opus 4.8 · GPT-5.5 · Claude Opus 4.7 · Opus 4.7 · Composer 2.5
May 27, 2026
Codex optimized its own AGENTS.md against real Stet repo tasks. The best candidate improved the training slice, then regressed enough on a clean holdout that it was not safe to ship.
Codex · GPT-5.5 · GPT-5.4
May 18, 2026
A fresh GPT-5.5 Codex high rerun on 21 clean GraphQL-go-tools tasks compared with the May 5 GPT-5.5 high run. The rerun was directionally worse on tests, equivalence, and review pass count, but the evidence is mixed and does not show a broad quality collapse.
GPT-5.5 · Codex
May 12, 2026
Claude Opus 4.7 reasoning-effort curve on 29 matched GraphQL-go-tools tasks: low, medium, high, xhigh, and max. Medium wins the behavioral metrics; more reasoning does not reliably buy better patches.
Claude Opus 4.7 · Opus 4.7
May 7, 2026
An interactive GPT-5.5 Codex reasoning-effort curve on 26 matched GraphQL-go-tools tasks: low, medium, high, and xhigh.
GPT-5.5 · Codex
May 1, 2026
Opus 4.7 vs GPT-5.5 vs GPT-5.4 on 56 real coding tasks across two open-source repos. Opus writes smaller patches; GPT-5.5 writes patches that more often survive review.
GPT-5.5 · GPT-5.4 · Claude Opus 4.7 · Opus 4.7
April 17, 2026
Three Opus snapshots, same 12/28 test pass rate. Above the gate, 4.7 is directionally better — more disciplined, not fundamentally smarter.
Claude Opus 4.7 · Opus 4.7 · Claude Opus 4.6 · Opus 4.6