Deep Dive · Disputed Benchmark Score

Claude Fable 5 on SWE-Bench Pro: Where the 80.3% Score Comes From, and Whether to Trust It

Claude Fable 5 scored 80.3% on SWE-Bench Pro -- one of the highest published numbers out there. But how was it actually measured, and why do independent evaluators report different results? Here's both sides.

#80.3% Score#SWE-Bench Pro#Methodology Dispute#Claude Fable 5
Read this first: the score is disputed

80.3% is Anthropic's official figure, produced with Anthropic's own evaluation scaffolding -- not independently reproduced by a neutral third-party test harness. Reports from independent evaluators, including techjacksolutions.com, note that reproduced scores differ from the official number. See the "Key Clarification" section below before drawing conclusions from the headline figure alone.

Last verified: 2026-08-08.

SWE-Bench Pro Leaderboard Comparison

The SWE-Bench Pro scores Anthropic published at Fable 5's launch, compared against other flagship models from the same period. All figures come from official/evaluator-published materials -- pay attention to the "Evaluator/Methodology" column, since these weren't all produced by the same neutral test run.

Model SWE-Bench Pro Score Evaluator / Methodology Note
Claude Fable 5 80.3% Anthropic's own evaluation scaffolding Tied with Mythos 5
Claude Opus 4.8 69.2% Anthropic's own evaluation scaffolding Previous-generation Claude flagship, included as a baseline
GPT-5.5 58.6% Comparison data published by Anthropic OpenAI's flagship model; not independently verified by OpenAI
Gemini 3.1 Pro 54.2% Comparison data published by Anthropic Google's flagship model; not independently verified by Google

A Closer Look at the Score

How the 80.3% was measured

SWE-Bench Pro is a benchmark built from real-world software engineering tasks -- models work inside codebases that resemble actual production repos, fixing bugs and making changes, at a level of complexity closer to real development work than earlier SWE-Bench versions. Claude Fable 5 scored 80.3% on this benchmark, tying with Mythos 5, launched the same day -- one of the headline numbers in Anthropic's official release materials. The score itself is real; it wasn't fabricated. The question isn't whether the number exists, it's under what conditions it was produced.

Key clarification: official evaluation methodology, not a neutral reproduction

The 80.3% figure was produced using Anthropic's own evaluation scaffolding -- in other words, this is an official-methodology score, not one independently reproduced by a neutral third-party test harness. That distinction matters: the scaffolding determines how the model reads each task, how it calls tools, whether it can retry after a failure, and how much context budget it gets -- details that materially affect the final pass rate. And when the vendor being tested builds and tunes that scaffolding itself, there's inherent room for it to be more favorable to their own model. techjacksolutions.com published a skeptical report specifically noting that when independent evaluators reproduced SWE-Bench Pro using their own methodology, the resulting scores differed from Anthropic's official number. This doesn't mean 80.3% is fake or meaningless -- it means it's a score under official evaluation methodology, which is a different thing from an independent reproduction, and readers deserve to know both sides rather than just the number from the launch announcement.

What this kind of score is -- and isn't -- good for

Benchmark scores are a useful data point, but not the whole story for model selection. SWE-Bench Pro reflects performance on a specific distribution of tasks -- bug-fix-style work in real-world codebases -- and your own use case might look very different: greenfield builds, complex system design, large-scale refactors under long context windows. The score doesn't necessarily transfer. On top of that, inconsistent methodology (official scaffolding vs. independent reproduction) means numbers on the same leaderboard weren't necessarily produced under identical, fair conditions. When actually choosing a model, look beyond the benchmark score at your specific task type, context-length needs, and the cost differences between models -- and ideally run a small-scale test with your own real tasks before deciding.

How This Page Relates to the Other Two

These three pages are complementary, not redundant -- each one covers a different question.

Want the full leaderboard?

For the complete SWE-Bench Pro rankings across all major models, see the "SWE-Bench Pro 2026 Leaderboard" page. This page goes deep on one model, Fable 5, and the story behind its score -- it doesn't re-list the full rankings here.

Want the broader methodology critique?

The general methodological limitations of benchmarks like SWE-Bench -- task-selection bias, scaffolding dependence, how representative they are of production work -- are covered in depth in "SWE-Bench and the Reality Gap in Production." This page relies on that page's conclusions rather than re-arguing them here.

Using Fable 5 on QCode

QCode already offers Claude Fable 5, accessible with a single key. For pricing details and comparisons against other models, see the pricing guide -- that's not covered here.

Timeline

The key date currently on record.

2026-06-09

Claude Fable 5 launches, 80.3% score published

Anthropic launches Claude Fable 5; official evaluation materials show an 80.3% score on SWE-Bench Pro, tied with Mythos 5, launched the same day.

Primary Sources

the-decoder.com's reporting on the Claude Fable 5 launch and its published benchmark figures; techjacksolutions.com's skeptical report on the gap between the official SWE-Bench Pro score and independently reproduced results; vellum.ai's SWE-Bench Pro benchmark breakdown; Anthropic's official release page.

FAQ

Is Claude Fable 5's 80.3% on SWE-Bench Pro the highest score out there?

Among the models in Anthropic's own published comparison, 80.3% is the highest, tied with Mythos 5. But keep in mind this is under Anthropic's own evaluation methodology -- independent reproductions by other organizations don't necessarily produce the same ranking. See the next question.

Is this score trustworthy? Why is it disputed?

The number itself is a real, officially published figure -- it wasn't made up. But it was produced using Anthropic's own evaluation scaffolding, an official-methodology result rather than an independent third-party reproduction. techjacksolutions.com's skeptical reporting notes that independent evaluators who reproduced SWE-Bench Pro got scores that differ from the official number. The dispute is about who built the test conditions and whether that setup inherently favors the model being tested -- not about whether the score exists at all.

What's the actual difference between official methodology and independent reproduction?

The evaluation scaffolding determines how a model reads task descriptions, calls tools, retries after failures, and how much context/step budget it's given -- implementation details that materially affect the final pass rate. When the vendor whose model is being tested builds and tunes that scaffolding themselves, results can diverge from what a neutral third party running all models under one consistent standard would get. This is a general issue with vendor-published benchmark numbers, not something unique to Fable 5.

How much should this score actually influence my model choice?

Treat it as one meaningful signal, not the only one. It reflects roughly how capable the model is on bug-fix-style tasks in real-world codebases, which may or may not match your own use case (task type, context length, cost budget). We'd suggest checking the full "SWE-Bench Pro 2026 Leaderboard" for the broader comparison, then validating against your own real tasks at small scale.

Create a QCode account

Plans from ¥60 / $8.57. One key across Claude, GPT, and Gemini models.

Related Reading

This site is not affiliated with Anthropic. SWE-Bench Pro scores follow official release materials and third-party reporting; information on this page is current as of 2026-08-08 and may change. Differences in evaluation methodology mean scores from different sources may not be directly comparable.