🏆 2026-08-18 board (aggregated methodology)

SWE-Bench Pro 2026 — The Realistic Benchmark for AI Coding Models

The 08-18 aggregated board: Mythos 5 80.3, Fable 5 80.0 and Opus 5 79.2 lead, while open-weight Qwen3.8 Max 67.7 overtakes GPT-5.6 Sol 64.6. Mind the methodology: these are aggregated figures — vendor self-reports and Scale's official board use different methodologies; see our 'Three First Places' page for the differences.

SWE-Bench Pro vs Verified — The Benchmark

SWE-Bench Verified (popular from 2024): ~500 human-verified GitHub issue-fix tasks. SWE-Bench Pro (mainstream from 2026): same approach but harder — longer context, more files touched, closer to real PR workflows. Verified plateaued near 80% (5.2-Codex), Pro still has ~40 points of headroom and is now the canonical benchmark. Terminal-Bench 2.0 measures terminal-agent tasks; OSWorld measures GUI-operating tasks; GDPval measures professional knowledge.

2026 Q1-Q2 Score Evolution

Early evolution: GPT-5.2-Codex (2026-01) Verified 80.0% / Pro 56.4%; GPT-5.3-Codex (2026-02) Pro 56.8% (this model was retired in 2026-08); GPT-5.5 Pro 58.6%. The 08-18 aggregated board (llm-stats / benchlm methodology): Mythos 5 80.3, Fable 5 80.0, Opus 5 79.2, Qwen3.8 Max 67.7, GPT-5.6 Sol 64.6, GLM-5.2 62.1. Note: Anthropic's system card self-reports Fable 5 at 80.3 (tied with Mythos), while the aggregated board records 80.0 — both numbers exist; the difference is the evaluation methodology.

Cross-Model Comparison by Scenario

By the 08-18 aggregated board: Mythos 5 (80.3) > Fable 5 (80.0) > Opus 5 (79.2) lead, and Qwen3.8 Max (67.7) overtakes GPT-5.6 Sol (64.6). Match models to scenarios: 1) long-horizon agent / terminal tasks → see the Terminal-Bench 2.1 board (another page on this site); 2) deep Python refactors → the Claude line (Opus 5 / Fable 5); 3) open weights + value → Qwen3.8 Max / GLM-5.2 / DeepSeek V4 Pro; 4) the OpenAI ecosystem → GPT-5.6 Sol. For the methodology mess, see 'SWE-bench Pro's Three First Places'.

Unified Access to All Major Models via QCode

QCode.cc provides a single transparent API gateway to all major coding models from inside China — Claude (Opus 4.7 / Sonnet), GPT (5.5 / Codex family), Gemini (3) and more. One subscription, usage-based billing. In Claude Code, Codex CLI, Cursor, Cline, Continue, switching models is just changing base URL and model id — configure once, run everywhere.

FAQ

Is 58.6% on SWE-Bench Pro high? How should I read it?

Pro is far harder than Verified: Verified is saturated (80%+), while on Pro the top vendor self-reports sit around 80 and Scale's official unified-framework numbers come in lower. Before reading any Pro score, ask three things: which board, what scaffold, which effort tier — details in 'SWE-bench Pro's Three First Places'.

Why does general-purpose GPT-5.5 beat coding-specialized 5.3-Codex on Pro?

That's a first-half-of-2026 picture. The top of the 08-18 aggregated board is now three Claude models (Mythos 5 / Fable 5 / Opus 5), and GPT-5.3-Codex was retired in 2026-08 (migrate to the three GPT-5.6 tiers). Since 2026-09-03 the top of OpenAI's line is GPT-6 Astra — but note that OpenAI published no SWE-bench Verified score for it on either the model page or the system card, so it has no comparable position on this board. Do not treat numbers circulating elsewhere as official.

Opus 4.7's Pro score isn't published — how do I evaluate it?

On this 2026-08-18 aggregated board Opus 4.7 has no vendor-reported Pro score, so don't rank it by score. Its slot moved on to Opus 5; as of 2026-09-29 the newest Opus from Anthropic is Opus 5.5 (released 2026-09-22, $4 / $20). See the tracker /claude-opus-5-2-status-tracker and the specs in /claude-opus-5-5-guide.

Access All Major Coding Models Through QCode

Mythos 5 is restricted-supply; Fable 5 / Opus 5 / the full GPT-5.6 line / GLM-5.3 / Kimi K3 / DeepSeek V4 Pro are all callable through QCode's unified channel, billed by usage.

Start Your QCode Plan

Create a QCode account

Plans from ¥60 / $8.57. One key across Claude, GPT, and Gemini models.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.