Deep Dive · Disputed Benchmark Score

Claude Fable 5 on SWE-Bench Pro: Where the 80.3% Score Comes From, and Whether to Trust It

Claude Fable 5 scored 80.3% on SWE-Bench Pro -- one of the highest published numbers out there. But how was it actually measured, and why do independent evaluators report different results? Here's both sides.

#80.3% Score#SWE-Bench Pro#Methodology Dispute#Claude Fable 5
Read this first: the score is disputed

80.3% is Anthropic's official-methodology result, produced with its own evaluation scaffolding. The aggregated boards (llm-stats / benchlm, snapshots 2026-08-18/21) record Fable 5 at 80.0%, with Mythos 5 leading at 80.3% — both numbers genuinely exist; what differs is the methodology. Also note that Scale's official unified-framework board is lower across the board. For the full three-methodology comparison, see 'SWE-bench Pro's Three First Places' on this site.

Last re-verified: 2026-08-22 (both methodologies shown side by side: official self-report 80.3 / aggregated board 80.0).

SWE-Bench Pro Leaderboard Comparison

The SWE-Bench Pro scores Anthropic published at Fable 5's launch, compared against other flagship models from the same period. All figures come from official/evaluator-published materials -- pay attention to the "Evaluator/Methodology" column, since these weren't all produced by the same neutral test run.

Model SWE-Bench Pro Score Evaluator / Methodology Note
Claude Fable 5 80.3% Anthropic's own evaluation scaffolding Tied with Mythos 5
Claude Opus 4.8 69.2% Anthropic's own evaluation scaffolding Previous-generation Claude flagship, included as a baseline
GPT-5.5 58.6% Comparison data published by Anthropic OpenAI's flagship model; not independently verified by OpenAI
Gemini 3.1 Pro 54.2% Comparison data published by Anthropic Google's flagship model; not independently verified by Google

A Closer Look at the Score

How the 80.3% was measured

SWE-Bench Pro is an evaluation suite built on real-world software engineering tasks: models must complete fixes, refactors and similar work inside codebases that resemble real development scenarios — closer to production complexity than earlier SWE-Bench versions. Claude Fable 5 scored 80.3% in Anthropic's official materials (tied with Mythos 5), while the aggregated boards (llm-stats / benchlm) record 80.0% for Fable 5 and 80.3% for Mythos 5 — both numbers genuinely exist; the question is not which one is fabricated, but under which methodology each was measured.

Key clarification: official evaluation methodology, not a neutral reproduction

The 80.3% figure was produced using Anthropic's own evaluation scaffolding -- in other words, this is an official-methodology score, not one independently reproduced by a neutral third-party test harness. That distinction matters: the scaffolding determines how the model reads each task, how it calls tools, whether it can retry after a failure, and how much context budget it gets -- details that materially affect the final pass rate. And when the vendor being tested builds and tunes that scaffolding itself, there's inherent room for it to be more favorable to their own model. techjacksolutions.com published a skeptical report specifically noting that when independent evaluators reproduced SWE-Bench Pro using their own methodology, the resulting scores differed from Anthropic's official number. This doesn't mean 80.3% is fake or meaningless -- it means it's a score under official evaluation methodology, which is a different thing from an independent reproduction, and readers deserve to know both sides rather than just the number from the launch announcement.

What this kind of score is -- and isn't -- good for

Benchmark scores are a useful data point, but not the whole story for model selection. SWE-Bench Pro reflects performance on a specific distribution of tasks -- bug-fix-style work in real-world codebases -- and your own use case might look very different: greenfield builds, complex system design, large-scale refactors under long context windows. The score doesn't necessarily transfer. On top of that, inconsistent methodology (official scaffolding vs. independent reproduction) means numbers on the same leaderboard weren't necessarily produced under identical, fair conditions. When actually choosing a model, look beyond the benchmark score at your specific task type, context-length needs, and the cost differences between models -- and ideally run a small-scale test with your own real tasks before deciding.

How This Page Relates to the Other Two

These three pages are complementary, not redundant -- each one covers a different question.

Want the full leaderboard?

For the complete SWE-Bench Pro rankings across all major models, see the "SWE-Bench Pro 2026 Leaderboard" page. This page goes deep on one model, Fable 5, and the story behind its score -- it doesn't re-list the full rankings here.

Want the broader methodology critique?

The general methodological limitations of benchmarks like SWE-Bench -- task-selection bias, scaffolding dependence, how representative they are of production work -- are covered in depth in "SWE-Bench and the Reality Gap in Production." This page relies on that page's conclusions rather than re-arguing them here.

Using Fable 5 on QCode

QCode already offers Claude Fable 5, accessible with a single key. For pricing details and comparisons against other models, see the pricing guide -- that's not covered here.

Timeline

The key date currently on record.

2026-06-09

Claude Fable 5 launches, 80.3% score published

Anthropic launches Claude Fable 5; official evaluation materials show an 80.3% score on SWE-Bench Pro, tied with Mythos 5, launched the same day.

Primary Sources

the-decoder.com's reporting on the Claude Fable 5 launch and its published benchmark figures; techjacksolutions.com's skeptical report on the gap between the official SWE-Bench Pro score and independently reproduced results; vellum.ai's SWE-Bench Pro benchmark breakdown; Anthropic's official release page.

FAQ

Is Claude Fable 5's 80.3% on SWE-Bench Pro the highest score out there?

It depends on the methodology: in Anthropic's official materials Fable 5 and Mythos 5 tie at 80.3%; on the aggregated boards (llm-stats / benchlm, 08-18 snapshot) Mythos 5 leads at 80.3%, Fable 5 is second at 80.0%, and Opus 5 third at 79.2%. Please cite the methodology when quoting.

Is this score trustworthy? Why is it disputed?

The score itself is a genuinely published number, not a fabrication; but 80.3% is the official-methodology result from Anthropic's own evaluation scaffolding, while the aggregated boards (llm-stats / benchlm) record 80.0%. Scale's official board on the unified SEAL framework is lower still across the board. The essence of the dispute is 'who built the test conditions and which data split was used', not whether the score exists. For the three-methodology comparison, see 'SWE-bench Pro's Three First Places'.

What's the actual difference between official methodology and independent reproduction?

The evaluation scaffolding determines how a model reads task descriptions, calls tools, retries after failures, and how much context/step budget it's given -- implementation details that materially affect the final pass rate. When the vendor whose model is being tested builds and tunes that scaffolding themselves, results can diverge from what a neutral third party running all models under one consistent standard would get. This is a general issue with vendor-published benchmark numbers, not something unique to Fable 5.

How much should this score actually influence my model choice?

Treat it as one meaningful signal, not the only one. It reflects roughly how capable the model is on bug-fix-style tasks in real-world codebases, which may or may not match your own use case (task type, context length, cost budget). We'd suggest checking the full "SWE-Bench Pro 2026 Leaderboard" for the broader comparison, then validating against your own real tasks at small scale.

Create a QCode account

Plans from ¥60 / $8.57. One key across Claude, GPT, and Gemini models.

Related Reading

This site is not affiliated with Anthropic. SWE-Bench Pro scores follow official release materials and third-party reporting; information on this page is current as of 2026-08-08 and may change. Differences in evaluation methodology mean scores from different sources may not be directly comparable.