Benchmarks and leaderboards
SWE-bench, Terminal-Bench and other public results — and what they actually tell you.
Terminal-Bench 3.0 Explained
TB 3.0 vs. 2.1, current rankings, and which leaders are sellable
Updated 2026-08-27
The Artificial Analysis Intelligence Index Explained
A full breakdown of the 9 weighted evaluations and common misreadings
Updated 2026-08-27
Terminal-Bench 2.1 Leaderboard Explained
Three boards, three winners: the official 83.8 / AA 89.5 / vals 85.77 gap fully explained
Updated 2026-08-22
SWE-bench Pro's Three First Places Explained
80.3 / 80.0 / the official board: the truth about data splits and scaffolds
Updated 2026-08-22
Evaluating Ornith-1.5 yourself
Judge a new model against your own work instead of public leaderboards.
Updated 2026-08-20
Claude Fable 5 SWE-Bench Pro Score, Explained
Where the 80.3% figure comes from, official methodology vs. independent reproduction, and how much it should weigh in your decision
Updated 2026-08-08
From SWE-Bench to Production: How Benchmarks Do (Not) Translate to Delivery Speed in 2026
88% on SWE-Bench sounds great, but often only 30% on your private repos. Understand the gap and build your own verification harness.
Updated 2026-07-13
SWE-Bench Pro 2026 Coding Model Ranking: Opus 4.7 vs GPT-5.5 vs GPT-5.3-Codex
SWE-Bench Pro 2026: GPT-5.3-Codex 56.8% SOTA, GPT-5.5 58.6%, GPT-5.2 80% Verified, Opus 4.8 Fast Mode default. Scenario-based model selection guide.
Updated 2026-05-16