# Stet Stet measures AI coding agents on real repository work. It helps teams compare model releases, agent harnesses, shared instructions such as AGENTS.md, reasoning settings, and private rollout changes before they become defaults. ## Core Pages - [Home](https://www.stet.sh/) - [Private evals](https://www.stet.sh/private) - [Methodology](https://www.stet.sh/methodology) - [Model comparisons](https://www.stet.sh/model-comparisons) - [Instruction evals](https://www.stet.sh/instruction-evals) - [Leaderboard](https://www.stet.sh/leaderboard) - [Blog](https://www.stet.sh/blog) ## High-Signal Evidence - [Codex AGENTS.md optimization holdout](https://www.stet.sh/blog/how-i-used-codex-to-improve-its-own-agents-md) - [GPT-5.5 vs GPT-5.4 vs Opus 4.7 evidence](https://www.stet.sh/benchmarks/gpt-55-vs-opus-47) - [GPT-5.5 reasoning curve evidence](https://www.stet.sh/benchmarks/gpt-55-codex-graphql-reasoning-curve) - [Opus 4.7 reasoning curve evidence](https://www.stet.sh/benchmarks/opus-47-graphql-reasoning-curve) - [Opus 4.7 vs Opus 4.6 evidence](https://www.stet.sh/benchmarks/opus-4-7-vs-opus-4-6-zod) ## Primary Claims - Stet replays real merged-code tasks from repositories. - Stet scores tests plus quality dimensions such as equivalence, code review, footprint risk, runtime, and cost. - Public benchmark pages should be read with the stated methodology and caveats. - Private Stet runs apply the same measurement pattern to a user's own repo, agent stack, and standards.