Benchmark Explainer · Terminal-Bench 3.0

Terminal-Bench 3.0
Much harder than 2.1 — the leader has barely cracked 40%

TB 3.0 launched on 2026-08-19 targeting sub-30% first-attempt pass rates for baseline models; the current leader, Claude Opus 5, scores just 42.7%, while 2.1's leaderboard is saturated around 88%.

#TB 3.0#launched 2026-08-19#42.7% leader#GPU sandboxes

Four key facts

42.7%

Current leaderboard score

Claude Opus 5 leads TB 3.0 at 42.7% — currently the only model above 40%.

88.4%

TB 2.1's leader score (saturated)

TB 2.1's leader, Grok 4.6, scores 88.4%, with the top four all clustered in the 86%-89% range — very little separation left.

<30%

3.0's design target

The developers (Harbor, Laude Institute, Snorkel AI, Turing) explicitly targeted sub-30% first-attempt pass rates for baseline models.

7 domains

Domain coverage

3.0 spans 7 task domains, introducing GPU sandboxes (up to H100), multi-container topologies, live microservices, and stricter grading logic.

What actually changed between 3.0 and 2.1

TB 2.1's top models have generally crept up toward 90%, leaving too little separation to see real gaps between models — the classic 'benchmark saturation' problem. TB 3.0 doesn't just add more questions; it raises the bar at the infrastructure level: GPU sandboxes up to H100, multi-container topologies, integration with live-running microservices, and stricter 'if-and-only-if' grading logic (not fuzzy approximate scoring). The developers have stated publicly that 3.0's task difficulty targets sub-30% first-attempt pass rates for baseline models — which is exactly why leader Claude Opus 5 only scores 42.7%. Models didn't get worse; the ruler got stricter.

Latest development

TB 3.0 launched on 2026-08-19, with its first version (v0.1) covering 74 tasks across 7 domains. The current top 8 are: Claude Opus 5 (42.7%), GPT-5.6 Sol (34.6%), Claude Fable 5 (34.0%), GLM-5.3 (32.4%), Grok 4.6 (26.5%), Claude Opus 4.8 (21.1%), GPT-5.6 Terra (20.8%), and Cognition's SWE-1.7 Lightning (18.6%).

Timeline

2026-08-19

Terminal-Bench 3.0 launches; v0.1 covers 74 tasks across 7 domains.

Launch week

The first leaderboard is published, with Claude Opus 5 leading at 42.7% — currently the only model above 40%.

Ongoing

As an early version, the task set and leaderboard are expected to keep expanding; rankings will shift as the benchmark evolves.

Confirmed vs. needs ongoing attention

Confirmed

TB 3.0 launched 2026-08-19, jointly developed by Harbor, Laude Institute, Snorkel AI, and Turing; the current leader is Claude Opus 5 (42.7%). TB 2.1's leader Grok 4.6 sits at 88.4%, with the top four clustered in the 86%-89% range.

Needs ongoing attention

3.0 is still an early version (v0.1, 74 tasks); the task set and leaderboard are likely to keep expanding and shifting, so don't treat any single snapshot of rankings as a permanent conclusion.

TB 3.0 vs. TB 2.1

TB 2.1 (trending toward saturation)

Top models clustered at 86%-89%, with little separation left; still widely used, but it's getting hard to see real gaps between models.

TB 3.0 (pulling scores apart)

GPU sandboxes + multi-container + live microservices + strict grading; the leader has just cracked 40%, giving a better read on current models' real ceiling.

How to read this leaderboard

First check whether the domain coverage matches your actual use case (3.0 spans 7 domains, and not every task is relevant to what you do). Then look at the actual score gaps rather than rank order — TB 3.0's top few models already show meaningful separation (42.7% vs 34.6% vs 34.0%), more useful than 2.1's cluster within a few points of each other. The benchmark is still early, so revisit it periodically rather than trusting a single snapshot.

What to do on QCode

Of TB 3.0's current top 8, Claude Opus 5, GPT-5.6 Sol, Claude Fable 5, GLM-5.3, and Claude Opus 4.8 are all in QCode's sellable lineup — one key calls them at official price × service rate. Grok 4.6 and SWE-1.7 Lightning are not currently sold through QCode.

FAQ

Why are TB 3.0 scores so much lower than 2.1's across the board?

Models didn't get weaker — the grading got stricter. 3.0 introduces GPU sandboxes, multi-container topologies, live microservices, and stricter grading logic, and the developers explicitly targeted sub-30% first-attempt pass rates for baseline models, so scores dropping across the board is expected.

Is TB 2.1 still worth referencing?

Yes, but its top models have largely saturated (86%-89% range), leaving little separation. If you're just comparing a handful of top models against each other, 3.0 is more informative.

How many tasks does TB 3.0 currently cover?

The first version (v0.1) covers 74 tasks across 7 domains; as an early version, the task set is expected to keep expanding.

Grok 4.6 topped TB 2.1 — why does it rank lower on 3.0?

The task difficulty and infrastructure requirements differ substantially between versions; 3.0 emphasizes realistic, complex multi-container/microservice scenarios, and different models perform quite differently there than on 2.1's task set — a ranking shift is expected.

Can I call every model on this leaderboard through QCode?

No. Of TB 3.0's top 8, Grok 4.6 and SWE-1.7 Lightning aren't currently sold through QCode; the other 5 (Claude Opus 5, GPT-5.6 Sol, Claude Fable 5, GLM-5.3, Claude Opus 4.8) can all be called through one QCode key.

What does TB 3.0's 'if-and-only-if' grading mean?

It refers to stricter automated grading — a task is only marked complete if it precisely satisfies the preset conditions, rather than fuzzy approximate scoring, which reduces false positives where something 'looks right but isn't fully correct.'

Sources

TB 3.0's launch date (2026-08-19), developers (Harbor, Laude Institute, Snorkel AI, Turing), domain count, design target (sub-30% baseline pass rate), and current leaderboard scores come from Terminal-Bench 3.0's official leaderboard page and related tech media coverage; TB 2.1 scores come from the 2.1 version of the same leaderboard series. Checked 2026-08-27.

Beyond the leaderboard: one key tests the mainstream models

Claude Opus 5, GPT-5.6 Sol, Claude Fable 5, GLM-5.3 and other TB 3.0 leaders — all callable with one QCode key.

Related reading

The leaderboard data on this page comes from Terminal-Bench's official leaderboard page and is not an official QCode statement; specific scores are governed by the live leaderboard, and we don't maintain our own independent ranking.