Benchmark Explainer · Artificial Analysis Intelligence Index

How the Artificial Analysis Intelligence Index Is Calculated
A weighted average of 9 evaluations, with published weights

The AA Intelligence Index is a weighted average of 9 sub-evaluations: Agents at 34%, Coding and Scientific Reasoning at 24% each, General at 18% — a model scoring differently across leaderboards or over time is usually about index version or data-pull timing, not the model changing.

#AA Intelligence Index#v4.1#9 weighted evals#pass@1

Four key facts

34%

Agents carries the heaviest weight

Among the 9 sub-evaluations, Agents has the largest share, split into GDPval-AA v2 (20%) and τ³-Banking (14%).

24% + 24%

Coding and Scientific Reasoning tie for second

Coding is 24% (Terminal-Bench v2.1 at 16%, SciCode at 8%); Scientific Reasoning is also 24% (Humanity's Last Exam 12%, GPQA Diamond and CritPt at 6% each).

18%

General carries the lightest weight

General is 18% (AA-LCR at 6%, AA-Omniscience contributing 8% accuracy plus 4% non-hallucination).

pass@1

Scoring method

Most sub-evaluations use pass@1 (correct on the first attempt), aggregated across repeated runs.

The full table of 9 sub-evaluations and their weights

The AA Intelligence Index v4.1 is a weighted average of nine evaluations: Agents at 34% (GDPval-AA v2 at 20%, τ³-Banking at 14%); Coding at 24% (Terminal-Bench v2.1 at 16%, SciCode at 8%); Scientific Reasoning at 24% (Humanity's Last Exam at 12%, GPQA Diamond at 6%, CritPt at 6%); and General at 18% (AA-LCR at 6%, AA-Omniscience contributing 8% accuracy and 4% non-hallucination). Scoring mainly uses pass@1 — the model must get it right on the first try, aggregated across multiple repeated runs.

Why the same model's score looks different in different places

Two models with the same composite score don't necessarily have the same capability profile — they can be entirely different combinations of strengths that happen to sum to the same weighted total. On top of that, Artificial Analysis keeps revising the index version (currently v4.1, introduced June 2026) and the specific sub-evaluation composition, and different versions carry different weights; even at the same version, different data-pull dates capture model updates and shifting eval environments, so the same model's score can drift.

Timeline

Early versions

Artificial Analysis introduces the Intelligence Index as a weighted composite score for comparing model capability across the board.

2026-06

Index v4.1 is introduced, publicly specifying the nine sub-evaluations and their weights (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%).

Ongoing

The sub-evaluation composition and weights get adjusted across versions, so scores from different data-pull dates may not line up exactly.

Confirmed vs. common misreading

Confirmed

AA Intelligence Index v4.1's nine sub-evaluations and their weights (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%, and their respective sub-weights) are publicly documented on Artificial Analysis's official methodology page; scoring is primarily pass@1.

Common misreading

Many people treat the composite score as a single-dimension capability ranking and use it directly to claim 'this one beats that one' — but the composite is a weighted sum across nine heterogeneous evaluations. Two models with similar totals can have very different strength profiles across Agents, Coding, and Scientific Reasoning; looking only at the total is easy to misread.

Composite score vs. sub-scores

Just the composite score

Good for a quick rough tier of capability, but doesn't tell you specifically where a model is strong or weak.

Breaking it down by sub-evaluation

If your use case leans heavily toward Agents or Coding, going straight to the relevant sub-evaluation (like the Terminal-Bench v2.1 component) is more targeted than the composite.

How to actually use this index

Step 1: figure out which of the nine categories (Agents, Coding, Scientific Reasoning, or General) is closest to your actual use case, and prioritize the relevant sub-score over the composite. Step 2: when comparing models, confirm both scores come from the same index version and the same data pull — cross-version or cross-date comparisons are easy to misread. Step 3: treat the composite as a rough filter, not a final answer — pair it with a small-scale test on your actual task before deciding.

What to do on QCode

Whatever dimension of the index matters most for your use case, one QCode key connects you to several top-scoring model families at once, so you can run a small-scale comparison on your actual task instead of deciding purely off a single composite number.

FAQ

How exactly is the AA Intelligence Index calculated?

It's a weighted average of nine sub-evaluations: Agents 34% (GDPval-AA v2 20% + τ³-Banking 14%), Coding 24% (Terminal-Bench v2.1 16% + SciCode 8%), Scientific Reasoning 24% (HLE 12% + GPQA Diamond 6% + CritPt 6%), General 18% (AA-LCR 6% + AA-Omniscience accuracy 8% + non-hallucination 4%).

Why does the same model's score look different than what I saw before?

The most common reasons are an index version upgrade (weights or sub-evaluation composition changed) or a different data-pull date (the model itself may have been updated, or the eval environment adjusted).

Do two models with the same composite score really have equal capability?

Not necessarily. The composite is a weighted sum, so two models can each be strong and weak in different sub-areas and just happen to land on a similar total — their actual capability profiles can differ a lot.

What does pass@1 mean?

It means the model gets one shot — it only counts as a pass if it's correct on the first attempt; most sub-evaluations run multiple repeats and aggregate the pass@1 results into a score.

Terminal-Bench v2.1 is just part of the Coding sub-score — why do other pages on this site treat it as a standalone leaderboard?

Terminal-Bench is both an independent public leaderboard and a component within the AA Index's Coding category (16 of Coding's 24 points). The two framings aren't contradictory — they're just different levels of granularity.

What else should I watch for when picking a model based on this index?

Treat the index as a general reference, not the only input. If your task is specific, go straight to the relevant sub-evaluation's score, and pair it with a small-scale test before making a final call.

Sources

The AA Intelligence Index v4.1's nine sub-evaluations, their weights, the scoring method (pass@1), and the version's introduction date (June 2026) come from Artificial Analysis's official methodology page. Checked 2026-08-27.

Don't decide off a single composite number

One QCode key connects you to several top-scoring model families for a small-scale test on your actual task.

Related reading

This page's explanation of the Artificial Analysis Intelligence Index methodology is based on its official public documentation and is not an official QCode statement; specific weights and sub-evaluation composition are governed by Artificial Analysis's live official page.