Leaderboard read · 2026-08-22 snapshot

Terminal-Bench 2.1
Three number ones — which to trust

Same benchmark, but the official board, Artificial Analysis and vals.ai each crown a different winner. The gap isn't score-gaming — it's harness, effort tier and snapshot dates.

#terminal-bench-2.1#89 tasks#Three leaderboards#TB 3.0 is here

The three leaders

83.8

Official board · Claude Code × Fable 5

tbench.ai, agent + model combo entry, xhigh tier, 2026-06-07 snapshot, $552.67 per run.

89.5

Artificial Analysis · GPT-5.6 Sol

xhigh reasoning tier. Opus 5 follows at 89.1, Grok 4.6 is third at 88.4 (via The Batch, 08-21).

85.77

vals.ai · GPT-5.6 Sol

Unified Terminus 2 harness, pass@1. Opus 5 at 84.64, Kimi K3 at 80.90.

89

Tasks

Terminal tasks from model training to sysadmin work, graded by difficulty; at launch no model exceeded 50% on the hard tier.

What Terminal-Bench 2.1 is

Terminal-Bench is a community open-source terminal-agent benchmark: models must actually complete tasks inside a sandboxed terminal — for example, "train a fasttext model on the yelp data in data/, keep it under 150MB and hit 0.62 accuracy on a private test set." Version 2.1 has 89 tasks spanning model training, system administration and more, scored pass@1: a task only counts if all its bundled pytest checks pass. Terminal work is the home turf of agents like Claude Code, Codex and Cursor, which makes this board a hard metric for agentic coding.

Why it's flaring up again this month

Since 08-08 multiple outlets have cited Artificial Analysis numbers: GPT-5.6 Sol 89.5 vs Claude Opus 5 89.1 — a 0.4 gap at the top; on 08-21 The Batch reported Grok 4.6 entering the top three at 88.4. Meanwhile the official tbench.ai leader is Claude Code × Fable 5 at 83.8 — both sets of numbers are right, they just don't measure the same kind of "model." And GLM-5.3's launch notes already reference Terminal-Bench 3.0, turning 2.1 into last generation's board.

Timeline

2026-06-07

Current top entry on the official board: Claude Code × Fable 5 (xhigh) at 83.8.

2026-08-08

The "Sol 89.5 vs Opus 5 89.1" Artificial Analysis numbers start spreading through media coverage.

2026-08-21

The Batch: Grok 4.6 at 88.4 joins the top three; the same-period vals.ai snapshot reads Sol 85.77 / Opus 5 84.64 / K3 80.90.

Confirmed vs easily misread

Confirmed

All three leaderboards genuinely exist with verifiable numbers: official tbench.ai (agent + model combos, harbor framework), vals.ai (unified Terminus 2 harness, pass@1), Artificial Analysis (direct model tests with labeled effort tiers). The differences come from harness and tier, not fabricated data.

Easily misread

Headlines like "model X tops Terminal-Bench" almost never say which board, which agent or which effort tier. The official board measures "agent + model" combos (e.g. Claude Code × Fable 5), not bare models; reading combo scores as model scores is the most common mistake.

Compare cost before scores

Scores are close

The top two on the official board differ by 0.7 (83.8 vs 83.1); on Artificial Analysis by 0.4. 2.1 is close to saturated for frontier models.

Costs differ several-fold

Per-run cost on the official board: $552.67 for the Fable 5 combo, $2,059.19 for the GPT-5.5 combo. Same score band, nearly 4x the budget.

How to rerun it yourself

The official framework is harbor: harbor run -d terminal-bench/terminal-bench-2-1 -a "<agent>" -m "<model>" -k 5. To compare models, fix the agent and effort tier; to pick an agent, fix the model. vals.ai and Artificial Analysis snapshots work as cross-checks.

On QCode

The models at the front of all three boards — Fable 5, Opus 5, Sonnet 5, GPT-5.5, the full GPT-5.6 line, Kimi K3, GLM-5.2 — can all be rerun one by one with the same key on QCode, at official pricing times our service rate, for far less than topping up each platform separately. The Grok series isn't connected on QCode; use official channels for those.

FAQ

So who actually leads Terminal-Bench 2.1?

Depends on the board: official tbench.ai says Claude Code × Fable 5 (83.8), Artificial Analysis says GPT-5.6 Sol (89.5), and the vals.ai snapshot also has Sol (85.77). Their harnesses and scoring differ.

Why are the official numbers so much lower?

The official board tests "agent + model" combos with per-entry snapshot dates; vals.ai retests under a unified Terminus 2 harness, and Artificial Analysis tests models directly with labeled effort tiers. Stronger harness, higher tier, higher score.

What do the 89 tasks cover?

Real tasks in a sandboxed terminal: model training, system administration, data processing and more, graded by difficulty. Scored pass@1 — all bundled pytest checks must pass.

Is Terminal-Bench 3.0 out?

Vendors already cite it: GLM-5.3's launch notes reference TB 3.0, and third-party comparison pages show Opus 5 at 42.7 / Sol at 34.6. It's significantly harder; 2.1's top scores are near saturation.

How much does a rerun cost?

Official board entries carry a cost column: $552.67 per run for the Fable 5 combo, $2,059.19 for the GPT-5.5 combo. A unified pay-as-you-go API platform is far cheaper than subscribing to each platform separately.

Which board should drive selection?

Pick agents on the official board (combo scores), bare models on vals.ai (unified harness), and use Artificial Analysis for quick cross-model reads. No single number should be the sole basis.

Sources

tbench.ai official board (fetched 2026-08-22), vals.ai Terminal-Bench 2.1 methodology and snapshots, Artificial Analysis (via felloai and The Batch, 2026-08-21), BuildApps (2026-08-08).

One key — run your own conclusion

Front-running models are callable on the same QCode endpoint at official pricing times our service rate; you control the rerun budget.