STET

Model comparisons

Stet compares coding models on real merged-code tasks, then scores the patches above test pass rate: equivalence, code review, footprint risk, runtime, and cost.

Public comparisons show how model and harness choices behaved on these repos. A private Stet run answers the same question with your PR history, tests, AGENTS.md, and review standards. Compare Claude Code vs Codex on your repo.

Latest evidence

Model studies

Current comparisons and reasoning studies, ordered by publication date. Each writeup links the decision back to task-level evidence and the limits of the local slice.

June 19, 2026

GLM 5.2 on 50 real Go and Rust PRs: last on quality, and not the cheapest

GLM 5.2 vs Composer 2.5 and the premium field on 50 real merged PRs from graphql-go-tools (Go) and sqlparser-rs (Rust). GLM lands last on craft and equivalence in both repos, costs about twice Composer, and writes more code than the human while missing the change. A routing guide for the new cheap model.

GLM 5.2 · Composer 2.5 · Opus 4.8 · GPT-5.5

June 14, 2026

When Fable 5 Is Worth the Premium

Fable 5 vs Opus 4.8, GPT-5.5 and two more on 30 real GraphQL Go Tools and SQLParser Rust tasks. The useful read is taskwise W/D/L: which model won which metric, on which task, and why.

Fable 5 · Opus 4.8 · Opus 4.7 · GPT-5.5 · Composer 2.5

May 18, 2026

GPT-5.5 High Regression Check on GraphQL-go-tools

A fresh GPT-5.5 Codex high rerun on 21 clean GraphQL-go-tools tasks compared with the May 5 GPT-5.5 high run. The rerun was directionally worse on tests, equivalence, and review pass count, but the evidence is mixed and does not show a broad quality collapse.

GPT-5.5 · Codex

April 17, 2026

Opus 4.7 vs Old Opus 4.6 vs New Opus 4.6

Three Opus snapshots, same 12/28 test pass rate. Above the gate, 4.7 is directionally better — more disciplined, not fundamentally smarter.

Claude Opus 4.7 · Opus 4.7 · Claude Opus 4.6 · Opus 4.6

Canonical comparison evidence

30 tasks · GraphQL Go Tools and SQLParser Rust

Fable 5 vs Opus 4.8, GPT-5.5, Opus 4.7, and Composer 2.5 routing comparison

Fable 5 led the cleaned n=30 local score, but the taskwise metric ledger points to routing: pay for Fable on hidden semantic and patch-shape risk, keep cheaper models in the loop for narrower work.

Winner
Fable led local score
Tradeoff
Opus 4.8 stayed the cheaper default; Fable was the escalation route for above-test risk.
Evidence table

56 tasks · Zod and graphql-go-tools

GPT-5.5 vs GPT-5.4 real coding benchmark

GPT-5.5 beat GPT-5.4 on tests, equivalence, review pass, and clean pass across 56 real coding tasks.

Winner
GPT-5.5
Tradeoff
GPT-5.4 was cheaper per task; GPT-5.5 produced far more clean passes.
Evidence table

26 tasks · graphql-go-tools

GPT-5.5 low vs medium vs high vs xhigh reasoning curve

GPT-5.5 Codex at four reasoning settings on 26 matched graphql-go-tools tasks: equivalence and review pass climbed sharply with reasoning, while tests were not monotonic.

Winner
High as default, xhigh for complex work
Tradeoff
Xhigh produced the best equivalence and review quality, but cost about 2.18x high per task.
Evidence table

GPT-5.4 vs Opus 4.7

GPT-5.4 was stronger on equivalence in the 56-task benchmark: 35 of 56 patches matched the human change, versus 19 of 56 for Opus 4.7. Opus had the lower footprint risk, 0.20 versus 0.34, which makes it the more conservative patch writer.

MetricOpus 4.7GPT-5.4
Tests pass33 / 5631 / 56
Equivalence19 / 5635 / 56
Clean pass10 / 5611 / 56
Footprint risk0.200.34