Terminal-Bench 2.1
Three number ones — which to trust
Same benchmark, but the official board, Artificial Analysis and vals.ai each crown a different winner. The gap isn't score-gaming — it's harness, effort tier and snapshot dates.
The three leaders
Official board · Claude Code × Fable 5
tbench.ai, agent + model combo entry, xhigh tier, 2026-06-07 snapshot, $552.67 per run.
Artificial Analysis · GPT-5.6 Sol
xhigh reasoning tier. Opus 5 follows at 89.1, Grok 4.6 is third at 88.4 (via The Batch, 08-21).
vals.ai · GPT-5.6 Sol
Unified Terminus 2 harness, pass@1. Opus 5 at 84.64, Kimi K3 at 80.90.
Tasks
Terminal tasks from model training to sysadmin work, graded by difficulty; at launch no model exceeded 50% on the hard tier.
What Terminal-Bench 2.1 is
Terminal-Bench is a community open-source terminal-agent benchmark: models must actually complete tasks inside a sandboxed terminal — for example, "train a fasttext model on the yelp data in data/, keep it under 150MB and hit 0.62 accuracy on a private test set." Version 2.1 has 89 tasks spanning model training, system administration and more, scored pass@1: a task only counts if all its bundled pytest checks pass. Terminal work is the home turf of agents like Claude Code, Codex and Cursor, which makes this board a hard metric for agentic coding.
Why it's flaring up again this month
Since 08-08 multiple outlets have cited Artificial Analysis numbers: GPT-5.6 Sol 89.5 vs Claude Opus 5 89.1 — a 0.4 gap at the top; on 08-21 The Batch reported Grok 4.6 entering the top three at 88.4. Meanwhile the official tbench.ai leader is Claude Code × Fable 5 at 83.8 — both sets of numbers are right, they just don't measure the same kind of "model." And GLM-5.3's launch notes already reference Terminal-Bench 3.0, turning 2.1 into last generation's board.
Timeline
Current top entry on the official board: Claude Code × Fable 5 (xhigh) at 83.8.
The "Sol 89.5 vs Opus 5 89.1" Artificial Analysis numbers start spreading through media coverage.
The Batch: Grok 4.6 at 88.4 joins the top three; the same-period vals.ai snapshot reads Sol 85.77 / Opus 5 84.64 / K3 80.90.
Confirmed vs easily misread
Confirmed
All three leaderboards genuinely exist with verifiable numbers: official tbench.ai (agent + model combos, harbor framework), vals.ai (unified Terminus 2 harness, pass@1), Artificial Analysis (direct model tests with labeled effort tiers). The differences come from harness and tier, not fabricated data.
Easily misread
Headlines like "model X tops Terminal-Bench" almost never say which board, which agent or which effort tier. The official board measures "agent + model" combos (e.g. Claude Code × Fable 5), not bare models; reading combo scores as model scores is the most common mistake.
Compare cost before scores
Scores are close
The top two on the official board differ by 0.7 (83.8 vs 83.1); on Artificial Analysis by 0.4. 2.1 is close to saturated for frontier models.
Costs differ several-fold
Per-run cost on the official board: $552.67 for the Fable 5 combo, $2,059.19 for the GPT-5.5 combo. Same score band, nearly 4x the budget.
How to rerun it yourself
The official framework is harbor: harbor run -d terminal-bench/terminal-bench-2-1 -a "<agent>" -m "<model>" -k 5. To compare models, fix the agent and effort tier; to pick an agent, fix the model. vals.ai and Artificial Analysis snapshots work as cross-checks.
On QCode
The models at the front of all three boards — Fable 5, Opus 5, Sonnet 5, GPT-5.5, the full GPT-5.6 line, Kimi K3, GLM-5.2 — can all be rerun one by one with the same key on QCode, at official pricing times our service rate, for far less than topping up each platform separately. The Grok series isn't connected on QCode; use official channels for those.
FAQ
So who actually leads Terminal-Bench 2.1?
Depends on the board: official tbench.ai says Claude Code × Fable 5 (83.8), Artificial Analysis says GPT-5.6 Sol (89.5), and the vals.ai snapshot also has Sol (85.77). Their harnesses and scoring differ.
Why are the official numbers so much lower?
The official board tests "agent + model" combos with per-entry snapshot dates; vals.ai retests under a unified Terminus 2 harness, and Artificial Analysis tests models directly with labeled effort tiers. Stronger harness, higher tier, higher score.
What do the 89 tasks cover?
Real tasks in a sandboxed terminal: model training, system administration, data processing and more, graded by difficulty. Scored pass@1 — all bundled pytest checks must pass.
Is Terminal-Bench 3.0 out?
Vendors already cite it: GLM-5.3's launch notes reference TB 3.0, and third-party comparison pages show Opus 5 at 42.7 / Sol at 34.6. It's significantly harder; 2.1's top scores are near saturation.
How much does a rerun cost?
Official board entries carry a cost column: $552.67 per run for the Fable 5 combo, $2,059.19 for the GPT-5.5 combo. A unified pay-as-you-go API platform is far cheaper than subscribing to each platform separately.
Which board should drive selection?
Pick agents on the official board (combo scores), bare models on vals.ai (unified harness), and use Artificial Analysis for quick cross-model reads. No single number should be the sole basis.
Sources
tbench.ai official board (fetched 2026-08-22), vals.ai Terminal-Bench 2.1 methodology and snapshots, Artificial Analysis (via felloai and The Batch, 2026-08-21), BuildApps (2026-08-08).
One key — run your own conclusion
Front-running models are callable on the same QCode endpoint at official pricing times our service rate; you control the rerun budget.
Related reading
SWE-bench Pro 2026 ranking
Latest standings on the long-horizon repo-level benchmark.
SWE-bench production reality
The gap between leaderboard scores and real engineering.
AI model radar 2026
Price and availability tracking across flagship families.
Leaderboard numbers are a 2026-08-22 snapshot and change as boards update. Not affiliated with any benchmark organization.