There is no single “AI model rankings” list, and there isn’t supposed to be one. Arena measures which model people prefer talking to, based on blind human votes. Artificial Analysis’s Intelligence Index measures composite benchmark accuracy across reasoning, knowledge, and professional tasks. Task-specific boards for coding or price-efficiency measure something narrower still. Worse, even the same benchmark’s same score for the same model varies significantly across sites that claim to mirror it β by roughly 20% within about two weeks, in one documented case below. The useful skill isn’t memorizing a single number; it’s knowing which ranking actually answers your question.
Two genuinely different kinds of ranking, by design
Arena and Artificial Analysis aren’t competing to produce the same answer β they’re measuring different things on purpose. As one leaderboard aggregator puts it plainly: “Arena leaderboards rank models by crowd-voted preference; this leaderboard ranks them by independent benchmark scores.” A model can lead one and lag the other simultaneously. As of mid-September 2026, Claude Fable 5.1 leads Artificial Analysis’s Intelligence Index, but Arena “has no votes for Fable 5.1 yet” β the newest top-ranked model by one measure hadn’t even accumulated enough head-to-head votes to appear on the other yet. Neither ranking is wrong; they’re answering different questions. If you want the model people find most pleasant or useful to talk to, a preference-vote leaderboard is the closer match. If you want the model that scores highest on structured reasoning and professional-task benchmarks, a composite index like Artificial Analysis’s is the closer match. Treating either one as “the” ranking misreads what it was built to measure.
The number that shouldn’t have moved
This is the part worth taking seriously before quoting any specific score from a roundup, including this site’s own past comparisons. Three sources, all claiming to report Claude Fable 5.1’s score on Artificial Analysis’s own Intelligence Index, published within about two weeks of each other in September 2026:
| Source | Date | Claude Fable 5.1’s reported AA Intelligence Index score |
|---|---|---|
| FelloAI’s monthly rankings roundup | ~September 4, 2026 | 66 β reported as #1 overall |
| ModelGrep’s leaderboard | ~September 15, 2026 | 53.4 β reported as #1 overall |
| BenchLM’s benchmark mirror | September 15, 2026 (“data verified”) | 53.4 β reported as third place, behind GPT-5.6 Sol’s claimed 58.9 |
Two of the three at least agree with each other on the raw number (53.4), even while disagreeing on who it puts in first place. The third, checked only about eleven days earlier, reports a score roughly 20% higher for the identical model on the identical named index. That gap is too large to be normal measurement noise for a static, already-released model β it’s far more likely that Artificial Analysis periodically renormalizes its own index as new models and benchmarks get added to the composite, and that third-party sites mirroring the index are working from whatever snapshot they last pulled. Either way, the practical lesson is the same: a specific benchmark score quoted in a roundup β any roundup, including one on this site β has a shelf life measured in days to a couple of weeks, not months. Check the primary source, Artificial Analysis’s own site, directly, and note the date you checked it, rather than trusting a number that’s already been copied through one or two intermediary sites.
The overall leader isn’t always the category leader
A single composite score also hides category-specific results that might matter more to you than the headline rank. Moonshot’s Kimi K3, an open-weight model, has been reported ranking third overall on Artificial Analysis’s Intelligence Index β ahead of every proprietary model except two β while separately taking the top spot on a coding-specific arena board. A business choosing a model purely off the top-line “#1 overall” figure would walk right past the model that’s actually strongest at the specific task it needs, simply because that strength doesn’t show up until you look at the task-specific board rather than the composite score.
What to actually check, by what you’re trying to decide
- General reasoning and professional-task capability β Artificial Analysis’s Intelligence Index is the closer fit, but check it directly on Artificial Analysis’s own site rather than a mirror, and note the date.
- Which model people actually enjoy using, or trust for open-ended conversation β a preference-vote leaderboard like Arena captures that better than a benchmark composite, precisely because it’s measuring human judgment rather than task accuracy.
- Coding specifically β a dedicated coding-arena or SWE-bench-style board beats a general composite; one coding-specific arena board currently shows a different leader than the general Intelligence Index does, which is the point. How you run the model matters too; see our guide to coding-agent harnesses for why the same model can score very differently depending on the surrounding tooling.
- Cost-efficiency β a price-adjusted value ranking, or better, running your own numbers through our AI Cost Calculator, since a model’s rank on a generic value index still assumes an average usage pattern that may not match yours.
- A specific head-to-head decision you’re actually facing β a dedicated comparison beats a leaderboard row every time, because it can show where the two disagree and why. Our GPT-6 Astra vs Claude comparison, our Muse Spark vs Gemini comparison, our DeepSeek V4 vs GLM-5.3 comparison, and our Qwen vs Claude Opus 5 comparison are all examples of exactly this β genuinely split verdicts that a single ranking number would have flattened into a false “winner.”
This is the same discipline covered from a different angle in our piece on why a static AI tools directory goes stale: a frozen number, whether it’s a price or a benchmark score, starts decaying the moment it’s published. A ranking is a snapshot of a specific measurement on a specific date β treat it as exactly that, not as a permanent verdict.
How to check a ranking yourself, in under a minute
Go to the primary tracker directly β Artificial Analysis’s own site for the Intelligence Index, Arena’s own site (arena.ai) for preference votes β rather than a roundup or a “mirror” site that claims to reproduce one. Note the date you checked, since the discrepancy documented above shows even a two-week-old number can already be meaningfully out of date. If a roundup or comparison cites a specific score without a check date attached, treat that number as approximate rather than exact.
What is the “correct” AI model ranking right now?
There isn’t one correct ranking β different leaderboards measure genuinely different things. A preference-vote board like Arena and a benchmark composite like Artificial Analysis’s Intelligence Index can legitimately disagree because they’re not trying to answer the same question.
Why do different sites report different scores for the same benchmark?
Benchmark composites like Artificial Analysis’s Intelligence Index get periodically renormalized as new models and evaluations are added, and third-party sites that mirror the index often work from an older snapshot. A documented case in September 2026 showed the same model’s reported score on the same named index varying by roughly 20% across sources checked about two weeks apart.
What’s the difference between Arena’s leaderboard and Artificial Analysis’s Intelligence Index?
Arena ranks models by blind human preference votes β which response people actually liked better in a head-to-head test. Artificial Analysis’s Intelligence Index ranks models by a composite of independent benchmark scores covering reasoning, knowledge, and professional tasks. They measure different things and can rank the same models differently.
Does the #1 overall model mean it’s the best choice for me?
Not necessarily. A composite “overall” score can hide category-specific leaders β a model ranked third overall has been documented topping a coding-specific board ahead of models that beat it on the general composite. Check the category that matches your actual task rather than the headline number.
How often do AI model rankings actually change?
Often enough that a specific score is only reliable for a couple of weeks at most. New model releases reshuffle composite indices regularly, and the indices themselves get periodically renormalized, which can shift a model’s reported score even without a new release.
Where should I check rankings myself instead of trusting a roundup?
Go directly to the primary source β Artificial Analysis’s own site for its Intelligence Index, or Arena’s own site for preference-vote rankings β rather than a secondary site that mirrors or aggregates the data, and note the date you checked.
Shurah is the founder of AI Tools Daily, tracking pricing, licensing and policy changes across AI tools so readers can make decisions without wading through marketing claims themselves.