STET

On 25 Stet Tasks, Opus 4.8 Wrote Smaller Patches. Opus 5 Searched Wider.

August 8, 2026 · Updated August 17, 2026

Opus 5 is the new cool kid on the block, beating Fable 5 in benchmarks, yet remaining strangely frustrating to work with in practice. In order to gain more insight into Opus 5's behavior and to see how it performed on my repo, I ran Opus 4.8 and Opus 5 on the same 25 tasks drawn from merged work in Stet's own repository. I ran each model once per task with medium reasoning and identical evaluation criteria.

TL;DR

  • The score tied: 9/25 strict test passes each: the same 8 tasks, plus one unique pass apiece.
  • Opus 5 searched wider and verified more. It used more shell commands on 18 of 25 tasks, more test commands on 15, and performed more revision passes on the files it touched.
  • Opus 4.8 stayed contained. It had a smaller patch footprint on 20 of 25 tasks, meaning it stayed closer to the change that was actually merged.
  • Costs landed in the same range: Opus 5 was ~1.4% cheaper on the typical task, with ~4% more tokens and ~4% longer wall-clock.

At a high level, the results look the same: both models passed 9 tasks. But within these passes, neither the patches nor the process to get there looked the same.

Opus 4.8 had a lower task footprint (measure of how much code changed compared to the merged change) on 20 of 25 tasks. Opus 5 ran more shell commands on 18, more test commands on 15, and touched more files on 12 while tying on 11. Total tool calls split almost evenly, 13 to 11 with one exact tie. The models spent nearly identical interaction budgets on opposite parts of the work: Opus 4.8 spent its budget on the edit; Opus 5 spent its budget discovering what to edit and how to validate that change.

This difference is why it's important to look beyond top-level pass rates. A test pass rate simply tells you whether the test suite accepted the final patch. It notably does not tell you how the agent searched, what it chose to verify, how much code it left for review, whether it ever reached the file that owned the requested behavior, or how maintainable the code it wrote is.

A test fail can also hide a materially correct patch that still behaves as intended. So, Stet runs a second check called equivalence, asking whether the agent patch made the same behavioral change as the merged patch, even when the underlying implementation differs. Equivalence moves both models the same way. Opus 4.8 was judged equivalent on 12 of 25 tasks and Opus 5 on 11, with both equivalent on 10: the 8 shared test passes plus 2 shared test failures (stet-66762b4c, stet-f0caded9) where both patches implemented the merged behavior but still missed something needed for the tests to pass. Under either lens, the models stay effectively tied.

Note: this is 25 matched tasks from one repository. What follows is a behavioral read of a few tasks, not a definitive model ranking.

Grading

The deterministic testing signal and the grader signals point in different directions. Footprint risk separates the two models cleanly: 20 of 25 pairs for Opus 4.8. Double-clicking the chart shows that the extra code is not coming just from the volume of additional tests: 40% of Opus 5's churn lands in test and fixture code versus 34% for Opus 4.8, and its non-test churn is still 1.46× larger. When our graders do pick up signal, they lean towards Opus 5 on the coherence, instruction adherence, edge-case handling, and maintainability dimensions.

paired grader signals · two lenses

Footprint separates the models cleanly, but the graders mostly don't

One deterministic artifact measure and one v1 judge stack over the full matched cohort. The leans that do exist point toward Opus 5 on judged quality, while the patch surface points the other way.
Bars show the share of matched pairs favoring each arm: Opus 4.8 left, Opus 5 right, draws at center.
Opus 4.8 favoreddraw bandOpus 5 favored

artifact + craft · n=25 · 0.25-point draw band

code review · n=24 · 0.25-point draw band

Hover or focus a row for counts and context.

Craft rows pair v1 grader scores (0–4, sonnet-4-6 judge) on all 25 matched tasks with a 0.25-point draw band; code-review rows use the same scale and band on 23 complete pairs — one Opus 4.8 patch lacked a complete review. Footprint risk is computed from the patch, not graded.

Looking at this data, we can put together a coherent hypothesis about what wider search and heavier test execution buy in practice: judged patch quality tilts slightly upward while the artifact surface tilts sharply upward. At this sample size, both signals are directional.

25 matched repository tasks · every dot is one task

The score tied, but footprint, time, and cost didn't

The headline starts with the strict score, then keeps grading evidence, time, and cost in view. The per-task distributions explain how activity and resource tails produce those counts.
Vertical ticks mark the arm median. Log scales keep small and large tasks legible together.

strict score

9 / 25 each

same 8 passes · 1 Opus 4.8-only, 1 Opus 5-only

grading evidence

Opus 4.8 lower20 tasks

Opus 5 lower5 tasks

footprint risk · lower is better

time

Opus 4.8 faster17 tasks

Opus 5 faster8 tasks

agent duration · lower is faster

cost

Opus 4.8 cheaper10 tasks

Opus 5 cheaper15 tasks

cache-aware cost · typical task −1.4%

footprint risk

Opus 4.8 lower20tasksOpus 5 lower5tasks

lower is better · linear

00.20.40.6Opus 4.8stet-15439c21 · Opus 4.8 · 0.041stet-232b50ba · Opus 4.8 · 0.325stet-2450ca2d · Opus 4.8 · 0.407stet-3169aefa · Opus 4.8 · 0.332stet-43493c52 · Opus 4.8 · 0.357stet-5276aa3f · Opus 4.8 · 0.313stet-66762b4c · Opus 4.8 · 0.206stet-6f84e978 · Opus 4.8 · 0.447stet-89dfbc27 · Opus 4.8 · 0.010stet-b28f4fe7 · Opus 4.8 · 0.011stet-207089c5 · Opus 4.8 · 0.372stet-2826765e · Opus 4.8 · 0.194stet-30506f8e · Opus 4.8 · 0.346stet-342d0ac0 · Opus 4.8 · 0.170stet-5d766dde · Opus 4.8 · 0.317stet-5e075641 · Opus 4.8 · 0.426stet-60a0e6c2 · Opus 4.8 · 0.189stet-87814583 · Opus 4.8 · 0.090stet-a8716b88 · Opus 4.8 · 0.391stet-bbbbae09 · Opus 4.8 · 0.205stet-d2ab0c25 · Opus 4.8 · 0.137stet-e41c2e2d · Opus 4.8 · 0.064stet-e928166f · Opus 4.8 · 0.268stet-f0caded9 · Opus 4.8 · 0.174stet-f8ae1177 · Opus 4.8 · 0.194Opus 5stet-15439c21 · Opus 5 · 0.020stet-232b50ba · Opus 5 · 0.459stet-2450ca2d · Opus 5 · 0.333stet-3169aefa · Opus 5 · 0.323stet-43493c52 · Opus 5 · 0.356stet-5276aa3f · Opus 5 · 0.322stet-66762b4c · Opus 5 · 0.346stet-6f84e978 · Opus 5 · 0.453stet-89dfbc27 · Opus 5 · 0.359stet-b28f4fe7 · Opus 5 · 0.010stet-207089c5 · Opus 5 · 0.407stet-2826765e · Opus 5 · 0.494stet-30506f8e · Opus 5 · 0.550stet-342d0ac0 · Opus 5 · 0.343stet-5d766dde · Opus 5 · 0.458stet-5e075641 · Opus 5 · 0.693stet-60a0e6c2 · Opus 5 · 0.268stet-87814583 · Opus 5 · 0.448stet-a8716b88 · Opus 5 · 0.519stet-bbbbae09 · Opus 5 · 0.303stet-d2ab0c25 · Opus 5 · 0.324stet-e41c2e2d · Opus 5 · 0.221stet-e928166f · Opus 5 · 0.602stet-f0caded9 · Opus 5 · 0.253stet-f8ae1177 · Opus 5 · 0.363

agent duration

Opus 4.8 faster17tasksOpus 5 faster8tasks

lower is faster · log

5m15m45mOpus 4.8stet-15439c21 · Opus 4.8 · 16.7 minstet-232b50ba · Opus 4.8 · 19.6 minstet-2450ca2d · Opus 4.8 · 16.5 minstet-3169aefa · Opus 4.8 · 3.4 minstet-43493c52 · Opus 4.8 · 8.3 minstet-5276aa3f · Opus 4.8 · 7.5 minstet-66762b4c · Opus 4.8 · 27.3 minstet-6f84e978 · Opus 4.8 · 7.1 minstet-89dfbc27 · Opus 4.8 · 16.8 minstet-b28f4fe7 · Opus 4.8 · 19.1 minstet-207089c5 · Opus 4.8 · 31.1 minstet-2826765e · Opus 4.8 · 63.8 minstet-30506f8e · Opus 4.8 · 38.4 minstet-342d0ac0 · Opus 4.8 · 61.1 minstet-5d766dde · Opus 4.8 · 11.0 minstet-5e075641 · Opus 4.8 · 29.8 minstet-60a0e6c2 · Opus 4.8 · 22.7 minstet-87814583 · Opus 4.8 · 16.8 minstet-a8716b88 · Opus 4.8 · 38.6 minstet-bbbbae09 · Opus 4.8 · 29.5 minstet-d2ab0c25 · Opus 4.8 · 13.7 minstet-e41c2e2d · Opus 4.8 · 14.2 minstet-e928166f · Opus 4.8 · 25.3 minstet-f0caded9 · Opus 4.8 · 40.5 minstet-f8ae1177 · Opus 4.8 · 46.5 minOpus 5stet-15439c21 · Opus 5 · 5.3 minstet-232b50ba · Opus 5 · 13.4 minstet-2450ca2d · Opus 5 · 22.2 minstet-3169aefa · Opus 5 · 18.6 minstet-43493c52 · Opus 5 · 2.8 minstet-5276aa3f · Opus 5 · 9.7 minstet-66762b4c · Opus 5 · 37.2 minstet-6f84e978 · Opus 5 · 35.3 minstet-89dfbc27 · Opus 5 · 33.2 minstet-b28f4fe7 · Opus 5 · 20.7 minstet-207089c5 · Opus 5 · 31.3 minstet-2826765e · Opus 5 · 30.6 minstet-30506f8e · Opus 5 · 43.8 minstet-342d0ac0 · Opus 5 · 16.9 minstet-5d766dde · Opus 5 · 11.4 minstet-5e075641 · Opus 5 · 31.5 minstet-60a0e6c2 · Opus 5 · 25.2 minstet-87814583 · Opus 5 · 33.7 minstet-a8716b88 · Opus 5 · 43.5 minstet-bbbbae09 · Opus 5 · 19.1 minstet-d2ab0c25 · Opus 5 · 26.3 minstet-e41c2e2d · Opus 5 · 38.7 minstet-e928166f · Opus 5 · 25.6 minstet-f0caded9 · Opus 5 · 11.3 minstet-f8ae1177 · Opus 5 · 28.0 min

cache-aware cost

Opus 4.8 cheaper10tasksOpus 5 cheaper15tasks

lower is cheaper · log

$0.5$2$8Opus 4.8stet-15439c21 · Opus 4.8 · $0.96stet-232b50ba · Opus 4.8 · $0.78stet-2450ca2d · Opus 4.8 · $0.74stet-3169aefa · Opus 4.8 · $0.83stet-43493c52 · Opus 4.8 · $0.63stet-5276aa3f · Opus 4.8 · $0.95stet-66762b4c · Opus 4.8 · $1.63stet-6f84e978 · Opus 4.8 · $1.11stet-89dfbc27 · Opus 4.8 · $0.68stet-b28f4fe7 · Opus 4.8 · $1.01stet-207089c5 · Opus 4.8 · $10.75stet-2826765e · Opus 4.8 · $14.56stet-30506f8e · Opus 4.8 · $15.32stet-342d0ac0 · Opus 4.8 · $9.67stet-5d766dde · Opus 4.8 · $2.69stet-5e075641 · Opus 4.8 · $10.65stet-60a0e6c2 · Opus 4.8 · $10.66stet-87814583 · Opus 4.8 · $5.90stet-a8716b88 · Opus 4.8 · $21.48stet-bbbbae09 · Opus 4.8 · $9.99stet-d2ab0c25 · Opus 4.8 · $5.87stet-e41c2e2d · Opus 4.8 · $4.01stet-e928166f · Opus 4.8 · $2.26stet-f0caded9 · Opus 4.8 · $3.78stet-f8ae1177 · Opus 4.8 · $24.26Opus 5stet-15439c21 · Opus 5 · $0.42stet-232b50ba · Opus 5 · $0.99stet-2450ca2d · Opus 5 · $0.68stet-3169aefa · Opus 5 · $0.84stet-43493c52 · Opus 5 · $0.39stet-5276aa3f · Opus 5 · $0.85stet-66762b4c · Opus 5 · $1.42stet-6f84e978 · Opus 5 · $1.18stet-89dfbc27 · Opus 5 · $1.24stet-b28f4fe7 · Opus 5 · $0.54stet-207089c5 · Opus 5 · $9.98stet-2826765e · Opus 5 · $15.65stet-30506f8e · Opus 5 · $14.25stet-342d0ac0 · Opus 5 · $4.90stet-5d766dde · Opus 5 · $2.95stet-5e075641 · Opus 5 · $9.94stet-60a0e6c2 · Opus 5 · $6.88stet-87814583 · Opus 5 · $13.77stet-a8716b88 · Opus 5 · $11.37stet-bbbbae09 · Opus 5 · $7.38stet-d2ab0c25 · Opus 5 · $14.11stet-e41c2e2d · Opus 5 · $11.71stet-e928166f · Opus 5 · $5.32stet-f0caded9 · Opus 5 · $2.95stet-f8ae1177 · Opus 5 · $17.36

total tokens

Opus 4.8 fewer9tasksOpus 5 fewer16tasks

lower is fewer · log

1M10MOpus 4.8stet-15439c21 · Opus 4.8 · 1.2Mstet-232b50ba · Opus 4.8 · 766Kstet-2450ca2d · Opus 4.8 · 837Kstet-3169aefa · Opus 4.8 · 949Kstet-43493c52 · Opus 4.8 · 482Kstet-5276aa3f · Opus 4.8 · 928Kstet-66762b4c · Opus 4.8 · 1.9Mstet-6f84e978 · Opus 4.8 · 1.2Mstet-89dfbc27 · Opus 4.8 · 869Kstet-b28f4fe7 · Opus 4.8 · 1.5Mstet-207089c5 · Opus 4.8 · 14.8Mstet-2826765e · Opus 4.8 · 20.4Mstet-30506f8e · Opus 4.8 · 24.0Mstet-342d0ac0 · Opus 4.8 · 14.3Mstet-5d766dde · Opus 4.8 · 3.2Mstet-5e075641 · Opus 4.8 · 15.0Mstet-60a0e6c2 · Opus 4.8 · 15.5Mstet-87814583 · Opus 4.8 · 7.4Mstet-a8716b88 · Opus 4.8 · 34.8Mstet-bbbbae09 · Opus 4.8 · 12.8Mstet-d2ab0c25 · Opus 4.8 · 7.7Mstet-e41c2e2d · Opus 4.8 · 4.8Mstet-e928166f · Opus 4.8 · 1.9Mstet-f0caded9 · Opus 4.8 · 4.7Mstet-f8ae1177 · Opus 4.8 · 37.5MOpus 5stet-15439c21 · Opus 5 · 488Kstet-232b50ba · Opus 5 · 1.3Mstet-2450ca2d · Opus 5 · 818Kstet-3169aefa · Opus 5 · 939Kstet-43493c52 · Opus 5 · 328Kstet-5276aa3f · Opus 5 · 875Kstet-66762b4c · Opus 5 · 1.4Mstet-6f84e978 · Opus 5 · 1.5Mstet-89dfbc27 · Opus 5 · 1.5Mstet-b28f4fe7 · Opus 5 · 641Kstet-207089c5 · Opus 5 · 14.0Mstet-2826765e · Opus 5 · 22.5Mstet-30506f8e · Opus 5 · 22.7Mstet-342d0ac0 · Opus 5 · 6.9Mstet-5d766dde · Opus 5 · 4.0Mstet-5e075641 · Opus 5 · 14.5Mstet-60a0e6c2 · Opus 5 · 9.9Mstet-87814583 · Opus 5 · 21.4Mstet-a8716b88 · Opus 5 · 17.4Mstet-bbbbae09 · Opus 5 · 9.7Mstet-d2ab0c25 · Opus 5 · 22.7Mstet-e41c2e2d · Opus 5 · 18.3Mstet-e928166f · Opus 5 · 7.7Mstet-f0caded9 · Opus 5 · 3.8Mstet-f8ae1177 · Opus 5 · 26.4M

shell calls

Opus 5 higher18tasksOpus 4.8 higher7tasks

higher is descriptive · linear

03570105140Opus 4.8stet-15439c21 · Opus 4.8 · 29stet-232b50ba · Opus 4.8 · 15stet-2450ca2d · Opus 4.8 · 14stet-3169aefa · Opus 4.8 · 12stet-43493c52 · Opus 4.8 · 5stet-5276aa3f · Opus 4.8 · 10stet-66762b4c · Opus 4.8 · 22stet-6f84e978 · Opus 4.8 · 11stet-89dfbc27 · Opus 4.8 · 12stet-b28f4fe7 · Opus 4.8 · 20stet-207089c5 · Opus 4.8 · 64stet-2826765e · Opus 4.8 · 37stet-30506f8e · Opus 4.8 · 68stet-342d0ac0 · Opus 4.8 · 60stet-5d766dde · Opus 4.8 · 19stet-5e075641 · Opus 4.8 · 80stet-60a0e6c2 · Opus 4.8 · 51stet-87814583 · Opus 4.8 · 31stet-a8716b88 · Opus 4.8 · 122stet-bbbbae09 · Opus 4.8 · 40stet-d2ab0c25 · Opus 4.8 · 21stet-e41c2e2d · Opus 4.8 · 26stet-e928166f · Opus 4.8 · 21stet-f0caded9 · Opus 4.8 · 28stet-f8ae1177 · Opus 4.8 · 75Opus 5stet-15439c21 · Opus 5 · 13stet-232b50ba · Opus 5 · 17stet-2450ca2d · Opus 5 · 20stet-3169aefa · Opus 5 · 14stet-43493c52 · Opus 5 · 3stet-5276aa3f · Opus 5 · 11stet-66762b4c · Opus 5 · 16stet-6f84e978 · Opus 5 · 23stet-89dfbc27 · Opus 5 · 17stet-b28f4fe7 · Opus 5 · 13stet-207089c5 · Opus 5 · 59stet-2826765e · Opus 5 · 73stet-30506f8e · Opus 5 · 120stet-342d0ac0 · Opus 5 · 54stet-5d766dde · Opus 5 · 41stet-5e075641 · Opus 5 · 89stet-60a0e6c2 · Opus 5 · 68stet-87814583 · Opus 5 · 84stet-a8716b88 · Opus 5 · 94stet-bbbbae09 · Opus 5 · 46stet-d2ab0c25 · Opus 5 · 134stet-e41c2e2d · Opus 5 · 80stet-e928166f · Opus 5 · 69stet-f0caded9 · Opus 5 · 32stet-f8ae1177 · Opus 5 · 103

test commands

Opus 5 higher15taskstied3tasksOpus 4.8 higher7tasks

higher is descriptive · linear

010203040Opus 4.8stet-15439c21 · Opus 4.8 · 6stet-232b50ba · Opus 4.8 · 2stet-2450ca2d · Opus 4.8 · 3stet-3169aefa · Opus 4.8 · 2stet-43493c52 · Opus 4.8 · 2stet-5276aa3f · Opus 4.8 · 3stet-66762b4c · Opus 4.8 · 6stet-6f84e978 · Opus 4.8 · 2stet-89dfbc27 · Opus 4.8 · 3stet-b28f4fe7 · Opus 4.8 · 3stet-207089c5 · Opus 4.8 · 17stet-2826765e · Opus 4.8 · 8stet-30506f8e · Opus 4.8 · 23stet-342d0ac0 · Opus 4.8 · 16stet-5d766dde · Opus 4.8 · 6stet-5e075641 · Opus 4.8 · 22stet-60a0e6c2 · Opus 4.8 · 8stet-87814583 · Opus 4.8 · 12stet-a8716b88 · Opus 4.8 · 18stet-bbbbae09 · Opus 4.8 · 9stet-d2ab0c25 · Opus 4.8 · 1stet-e41c2e2d · Opus 4.8 · 7stet-e928166f · Opus 4.8 · 3stet-f0caded9 · Opus 4.8 · 7stet-f8ae1177 · Opus 4.8 · 14Opus 5stet-15439c21 · Opus 5 · 0stet-232b50ba · Opus 5 · 2stet-2450ca2d · Opus 5 · 6stet-3169aefa · Opus 5 · 2stet-43493c52 · Opus 5 · 1stet-5276aa3f · Opus 5 · 4stet-66762b4c · Opus 5 · 4stet-6f84e978 · Opus 5 · 7stet-89dfbc27 · Opus 5 · 5stet-b28f4fe7 · Opus 5 · 2stet-207089c5 · Opus 5 · 11stet-2826765e · Opus 5 · 17stet-30506f8e · Opus 5 · 37stet-342d0ac0 · Opus 5 · 10stet-5d766dde · Opus 5 · 11stet-5e075641 · Opus 5 · 22stet-60a0e6c2 · Opus 5 · 13stet-87814583 · Opus 5 · 26stet-a8716b88 · Opus 5 · 13stet-bbbbae09 · Opus 5 · 10stet-d2ab0c25 · Opus 5 · 13stet-e41c2e2d · Opus 5 · 20stet-e928166f · Opus 5 · 15stet-f0caded9 · Opus 5 · 10stet-f8ae1177 · Opus 5 · 17

what the footprint is made of

patch churn = lines added + deleted, summed over 25 tasks

solid: non-test code · light: test and fixture code

Opus 4.87,223 + 3,733 · 34% test
Opus 510,513 + 7,062 · 40% test

Opus 5's total churn is 1.64× Opus 4.8's, and more of it is test code — but the gap is not only tests: non-test churn alone is still 1.46× larger.

Hover or focus a row for the arm median.

Per-task values from the matched arm summaries of the full n=25 cohort. Pairwise leans count which arm recorded the better value on each matched task; higher activity is descriptive, not better. Footprint composition sums the arm summaries' surface breakdown into non-test versus test and fixture churn. No recalibrated uncertainty interval exists for this derivation.

Every task, side by side

Aggregates hide individual anecdotes that are useful for understanding model behavior. Let's dive into a few!

25 matched repository tasks · footprint risk per task

Which tasks drive the 20-to-5 footprint split

Opus 4.8 in grey, Opus 5 in vermillion, sorted by the gap between them. Squares at the right mark the strict result for each arm: filled for pass, hollow for fail.
Each row is one clean matched task. Lower footprint risk is better.
Opus 4.8Opus 5dashed: arm means 0.24 / 0.37
0.00.20.40.68781458389dfbc27e928166f2826765e5e07564130506f8ed2ab0c25342d0ac0f8ae1177e41c2e2d5d766dde66762b4c232b50baa8716b88bbbbae0960a0e6c2f0caded9207089c55276aa3f6f84e978b28f4fe743493c523169aefa15439c212450ca2d

Hover or focus a row for exact values.

Footprint risk is a deterministic patch-surface measure; lower is better. Dashed lines mark the arm means. Full n=25 matched descriptive cohort; per-task values come from the retained matched arm summaries.

Three things stand out. The split here is broad, not just driven by a few outliers: Opus 4.8's footprint advantage spans both small fixes and large multi-file changes. On the five widest tasks, the footprint gap alone exceeds Opus 4.8's average footprint across the whole cohort (0.24). Among the five tasks where Opus 5 left the smaller footprint, the most consequential is stet-2450ca2d, the only task Opus 5 passed while Opus 4.8 failed.

Opus 4.8 stayed closer to the patch it first understood

Footprint risk is Stet's deterministic measure of patch surface: files touched, churn, size, and overlap with the merged diff. A lower footprint score means that the agent's patch is more similar to what was merged previously. It says nothing about correctness, only surface.

stet-89dfbc27 shows why containment can be valuable. The task was to restore ignored files to Stet's synthetic base commit. Both agents found the production fix: add --force to git add -A.

Opus 4.8 changed one production file, added no test, and passed. Opus 5 made the same production change and then added a 141-line end-to-end test. Its test compiled and exercised a real boundary. It also turned a small repair into a much larger surface. Opus 5 spent nearly three times as long and 83% more recorded cost to produce the same accepted implementation plus broader verification.

stet-2450ca2d required two new test-file patterns in internal/gitops/testclassifier.go. Opus 4.8 edited internal/validate/footprint_risk.go, an adjacent consumer of the classifier output. It tested the function it changed, but never reached the owner of the requested behavior. Opus 5 found testclassifier.go, added both patterns, and passed strict and equivalence evaluations. Equivalence asks whether the agent patch made the same material behavioral change as the merged patch, rather than only whether deterministic testing passed.

Opus 4.8's patch was centered around the wrong owner. Note what else this task shows: it is one of only five where Opus 5 left the smaller footprint. When Opus 5's broader search finds the right owner, its wider exploration does not necessarily translate into a bigger patch.

A broader version of the same failure appears in stet-5d766dde. Opus 4.8 implemented the network-posture portion but omitted the paired schema/cache bump, mount stripping, and standalone resolver.

In summary, Opus 4.8's trajectory profile pays off when the task boundary is already known. It becomes more risky when the hard part is discovering how many owners the task actually has, and where that surface is, which is exactly the situation many large enterprise codebases find themselves in.

Opus 5 searched wider and kept working after the first edit

Looking at each task, Opus 5 recorded more shell calls on 18 pairs, more test commands on 15, and more distinct patched files on 12 with 11 ties. The patch-operations-per-file metric divides patch calls by distinct patched files on each task, so a higher value means more revisions per touched file, not better work.

Total tool calls are almost perfectly balanced between the two models. Opus 5 did not consume more interactions. It allocated more of them to the shell, test execution, and repeated editing.

That broader route is what passed stet-2450ca2d: six test commands instead of three, and the search continued past the adjacent consumer to the owning classifier. The implementation was small once the correct owner was found. The meat of the task was repository navigation to find the right surface.

The wider route created different failure modes on larger changes.

In stet-bbbbae09, Opus 5 recorded 24 patch calls across 8 files, renamed one required test, and omitted another. Opus 4.8 made 15 patch calls across 6 files and cleared the strict evaluator.

A longer trajectory is not waste, and a shorter one is not efficiency. Opus 5 often finished sooner and cheaper, yet missed named acceptance artifacts after more revisions. Opus 4.8 passed the evaluator, but its review artifact still raised API and authority concerns. Neither patch generalizes beyond its task.

stet-6f84e978 shows the valuable side of expansion. Opus 5 ran seven test commands against Opus 4.8's two and added a preservation test for an explicit non-Rust obligation. The stronger verification took 34.9 minutes instead of 6.1, while recorded cost rose only from $1.11 to $1.18. Wall time, tokens, cache mix, and price measure different parts of the trajectory.

Opus 5's wider search sometimes found the missing owner and sometimes created more room to drift from an exact contract. You can only see this when the comparison keeps the trajectory and the patch, not just the final test result.

selected task evidence · exact retained values

The headline changes when you inspect a task

These are the three task witnesses called out in the post. The bars are scaled within the selected task; the labels carry the exact values. Lower is better for footprint, time, tokens, and cost. Activity counts are descriptive.
Select a task to compare the same six measures behind the aggregate story.

stet-89dfbc27

same fix · broader verification

4.8 pass5 pass

Both agents found the required --force production fix. Opus 5 then added a 141-line end-to-end test, turning a small accepted repair into a much larger review surface.

footprint risk

lower is better

4.80.010
50.359

agent duration

lower is faster

4.816.8 min
533.2 min

total tokens

recorded input + output

4.8869K
51.48M

cache-aware cost

lower is cheaper

4.8$0.68
5$1.24

tool calls

higher is descriptive

4.816
523

test commands

higher is descriptive

4.83
55
Exact values come from the full n=25 matched arm summaries. The selected examples are illustrative witnesses, not a second sample or a model-wide verdict.

Time, tokens, and cost split in different directions

Three resource measurements answer three different questions. Agent duration is wall-clock time from the run's start to finish. Total tokens combine recorded input and output, including cached input. Cache-aware cost applies each model's price schedule to fresh input, cached input, and output. Opus 4.8 finished sooner on 17 pairs, Opus 5 cost less on 15, and the typical-task cost estimate landed just below Opus 4.8 at −1.4%.

Opus 5 used fewer tokens on 16 of 25 pairs and cost less on 15, so the counts lean its way. The paired-geometric magnitude points the other way on tokens: on the pairs where Opus 5 used more, it used enough more to put its typical task token estimate 4.3% above Opus 4.8, while cost settled 1.4% below and duration ran 3.7% longer. The count says how often a direction occurred; the paired estimate says how large the typical change was with every task weighted equally.

Two shared passes show how wide the range is:

  • On stet-15439c21, Opus 5 finished a small deletion in 294 seconds, 488K tokens, and $0.42 — 3.3 times faster with 2.4 times fewer tokens than Opus 4.8. Both passed.
  • On stet-89dfbc27, Opus 5 added a large end-to-end test and used 70% more tokens, 83% more cost, and 2.8 times the duration. Both passed.

The tails lean one way. On four of 25 tasks, Opus 5 used more than 2.5 times Opus 4.8's tokens, peaking at 4.1 times on stet-e928166f. Opus 4.8's largest token excess in the other direction was 2.4 times. On stet-66762b4c, Opus 5 produced substantially more implementation and test code and ran 10 minutes longer, yet used fewer recorded tokens and dollars; both remained strict failures, and both cleared a separate adapted-reference check that tested the material behavior without revising the strict result.

There is no clean "faster model" or "cheaper model" in this cohort. Resource use follows what the agent decides to inspect, implement, and verify on each task.

What the eval doesn't see

The thing that seriously frustrates me (and everyone else I talk to) about Opus 5's day-to-day behavior is its extremely verbose, hard-to-parse prose, which doesn't appear in these numbers at all. This evaluation scores the artifact: the patch, the tests it ran, the trajectory of how the agent got there. It does not score the interaction with the agent that produced that result. Walls of explanation, the restated plans, the summaries of summaries, eyes glazing over, LGTM, ship it. None of the eight craft dimensions measures how much reading the human had to do to get the final patch.

Code-side verbosity, another noted issue with Opus, does actually show up in our footprint risk metric. Even so, Opus can be disciplined in its patches and still exhausting for interaction, and this evaluation is structurally blind to that. This is an artifact eval, not a collaboration eval.

The more agentic model

On these tasks, Opus 5 looks like the more agentic model. It performed broader searching of the repo to figure out the correct surface before committing to an edit, it went looking for the place that owned the behavior instead of patching the nearest consumer, and it decided to validate its own work, resulting in more test commands and more post-edit revisions, rather than stopping at the first patch that seemed right. It did all of that while staying in the same price range: cheaper on 15 of 25 tasks, about 1.4% cheaper on the typical one.

The cost of that behavior shows up in review surface rather than dollars: 20 of 25 tasks left a bigger patch that a human (supposedly) has to review. Opus 5 buys discovery and verification, and you pay in patch surface and a little wall-clock.

Despite the prickly personality, I'll be using Opus 5, or having Fable delegate to it, for my hardest and most demanding problems.

Again, this is an n=1 repository. Model choice is one harness lever alongside instruction files, skills, tools, and reasoning settings, and any of them can change how an agent searches, edits, tests, and stops. The decision belongs on your own merged work, where the task distribution represents your own challenges, and the code review costs are tangible.

Methodology

Every task is derived from work that was actually merged into Stet's own repository — a PR or commit, replayed from a frozen snapshot of the tree as it stood before that change, with the issue prompt and the evaluation commands carried along. Both models ran all 25 tasks in the same Claude Code harness, one attempt per model-task cell at medium reasoning, against identical evaluation criteria.

The pass/fail score counts a cell as a pass only when the selected tests accept the agent patch. The adaptive lower bound additionally counts a cell whose implementation diverged from the merged change but was judged to reach the same behavior. Equivalence asks whether the agent patch made the same material behavioral change as the merged human patch, rather than only whether commands passed. Footprint risk is deterministic, not judged: it scores patch surface from files touched, churn, size, and overlap with the merged change, and lower means less code for a reviewer to hold in their head. It says nothing about correctness. The eight craft dimensions and the code-review rubric are pointwise judge scores from claude-sonnet-4-6, paired per task under a 0.25-point draw band on the 0–4 scale.

FAQ

How did Opus 4.8 and Opus 5 score on these 25 matched tasks?

Each model recorded 9/25 strict functional passes and a separately reported 10/25 adaptive functional lower bound. Eight tasks passed for both, 15 failed for both, and each model had one unique strict pass.

How did the models behave differently?

Opus 4.8 had lower footprint risk on 20 of 25 pairs. Opus 5 recorded more shell calls on 18 pairs, more test commands on 15, and more patched files on 12 with 11 ties. These are local paired observations, not global model traits.

Was Opus 5 cheaper or faster?

Opus 5 cost less on 15 of 25 pairs and used fewer tokens on 16, but Opus 4.8 finished sooner on 17. The paired-geometric estimates put Opus 5 1.4% below on cost, 4.3% above on tokens, and 3.7% longer on duration; the canonical report carries no decision-grade uncertainty interval for these.

Which model should a team use?

This comparison supports a local routing hypothesis, not a universal default: test Opus 4.8 on contained work where the owner is known and patch surface is expensive, and test Opus 5 on discovery-heavy work with exact contract verification. Validate that split on your own merged changes.