Preference
7,999,020 votes across 400 models
Blind head-to-head votes. Measures what people prefer reading, which tracks style as well as correctness.
PublicAI Index
A single benchmark is easy to target and easy to overfit. This index normalizes recognised public leaderboards onto one scale, discounts each score by the uncertainty its publisher reports, shrinks thin evidence, and weights the rest by a scheme published beside the table — overall, by category, and by domain. Figures from launch posts and blogs are indexed too, marked ✱ and kept out of the headline. Every model opens into a card with how to call it, and agents can ask the same questions over MCP.
CitePublicAI Foundation (2026). PublicAI Index: a weighted aggregate of public model leaderboards, snapshot 2026-09-06. https://publicai.io/model-index
§ 1 · Ranking
63 ranked · 402 provisional · showing 100 of 443
| # | Model | Index | Sources | |
|---|---|---|---|---|
| 1 | Claude Fable 5.1Anthropic | 69.1 | LMAATBARCLB5/5 | Show details for Claude Fable 5.1 |
| 2 | GPT-6 AstraOpenAI | 67.4 | LMAATBARCLB4/5 | Show details for GPT-6 Astra |
| 3 | Claude Fable 5Anthropic | 66.6 | LMAATBARCLB5/5 | Show details for Claude Fable 5 |
| 4 | Claude Opus 5Anthropic | 65.8 | LMAATBARCLB5/5 | Show details for Claude Opus 5 |
| 5 | GPT-5.5OpenAI | 64.3 | LMAATBARCLB3/5 | Show details for GPT-5.5 |
| 6 | GPT-5.6 SolOpenAI | 64.1 | LMAATBARCLB5/5 | Show details for GPT-5.6 Sol |
| 7 | Kimi K3Moonshot | 62.0 | LMAATBARCLB4/5 | Show details for Kimi K3 |
| 8 | Muse Spark 1.3Meta | 61.9 | LMAATBARCLB2/5 | Show details for Muse Spark 1.3 |
| 9 | GPT-5.4OpenAI | 61.6 | LMAATBARCLB3/5 | Show details for GPT-5.4 |
| 10 | Muse Spark 1.2Meta | 60.3 | LMAATBARCLB3/5 | Show details for Muse Spark 1.2 |
| 11 | qwen3.8Alibaba | 60.0 | LMAATBARCLB3/5 | Show details for qwen3.8 |
| 12 | Claude Opus 4.6Anthropic | 59.7 | LMAATBARCLB3/5 | Show details for Claude Opus 4.6 |
| 13 | Claude Opus 4.7Anthropic | 59.7 | LMAATBARCLB2/5 | Show details for Claude Opus 4.7 |
| 14 | GLM-5.3Z.ai | 59.3 | LMAATBARCLB4/5 | Show details for GLM-5.3 |
| 15 | Gemini 3.5 FlashGoogle | 59.1 | LMAATBARCLB3/5 | Show details for Gemini 3.5 Flash |
| 16 | GPT-5.6 TerraOpenAI | 59.1 | LMAATBARCLB✱5/5 | Show details for GPT-5.6 Terra |
| 17 | Grok 4.6SpaceXAI | 58.9 | LMAATBARCLB5/5 | Show details for Grok 4.6 |
| 18 | Gemini 3.7 FlashGoogle | 58.6 | LMAATBARCLB5/5 | Show details for Gemini 3.7 Flash |
| 19 | Gemini 3.1 Pro PreviewGoogle | 58.3 | LMAATBARCLB4/5 | Show details for Gemini 3.1 Pro Preview |
| 20 | Muse Spark 1.1Meta | 58.2 | LMAATBARCLB2/5 | Show details for Muse Spark 1.1 |
| 21 | Claude Opus 4.8Anthropic | 58.0 | LMAATBARCLB4/5 | Show details for Claude Opus 4.8 |
| 22 | Gemini 3 ProGoogle | 57.7 | LMAATBARCLB2/5 | Show details for Gemini 3 Pro |
| 23 | Claude Sonnet 4.6Anthropic | 56.6 | LMAATBARCLB3/5 | Show details for Claude Sonnet 4.6 |
| 24 | Gemini 3.8 FlashGoogle | 56.3 | LMAATBARCLB4/5 | Show details for Gemini 3.8 Flash |
| 25 | GPT-5.2OpenAI | 56.1 | LMAATBARCLB3/5 | Show details for GPT-5.2 |
| 26 | Gemini 3.6 FlashGoogle | 56.0 | LMAATBARCLB4/5 | Show details for Gemini 3.6 Flash |
| 27 | DeepSeek V4 Pro 0813DeepSeek | 55.6 | LMAATBARCLB3/5 | Show details for DeepSeek V4 Pro 0813 |
| 28 | GLM-5.3 FlashZ.ai | 55.1 | LMAATBARCLB3/5 | Show details for GLM-5.3 Flash |
| 29 | Grok 4.5SpaceXAI | 55.0 | LMAATBARCLB5/5 | Show details for Grok 4.5 |
| 30 | Inkling SmallThinking Machines | 54.9 | LMAATBARCLB2/5 | Show details for Inkling Small |
| 31 | GPT-5.1OpenAI | 54.8 | LMAATBARCLB2/5 | Show details for GPT-5.1 |
| 32 | Qwen3.8 27BAlibaba | 54.7 | LMAATBARCLB✱3/5 | Show details for Qwen3.8 27B |
| 33 | Claude Sonnet 5Anthropic | 54.2 | LMAATBARCLB✱4/5 | Show details for Claude Sonnet 5 |
| 34 | Kimi K2.5Moonshot | 54.0 | LMAATBARCLB2/5 | Show details for Kimi K2.5 |
| 35 | GPT-5.6 LunaOpenAI | 53.9 | LMAATBARCLB✱5/5 | Show details for GPT-5.6 Luna |
| 36 | GLM-5.2Z.ai | 53.9 | LMAATBARCLB✱3/5 | Show details for GLM-5.2 |
| 37 | GLM-5Z.ai | 53.6 | LMAATBARCLB2/5 | Show details for GLM-5 |
| 38 | Mimo V2.5 ProXiaomi | 53.5 | LMAATBARCLB2/5 | Show details for Mimo V2.5 Pro |
| 39 | Gemini 2.5 ProGoogle | 53.0 | LMAATBARCLB2/5 | Show details for Gemini 2.5 Pro |
| 40 | DeepSeek V4 Flash 0731DeepSeek | 53.0 | LMAATBARCLB3/5 | Show details for DeepSeek V4 Flash 0731 |
| 41 | GPT-5OpenAI | 53.0 | LMAATBARCLB2/5 | Show details for GPT-5 |
| 42 | Kimi K2.6Moonshot | 52.6 | LMAATBARCLB2/5 | Show details for Kimi K2.6 |
| 43 | DeepSeek V3.2DeepSeek | 51.9 | LMAATBARCLB2/5 | Show details for DeepSeek V3.2 |
| 44 | Grok 4 Fast ReasoningSpaceXAI | 51.0 | LMAATBARCLB2/5 | Show details for Grok 4 Fast Reasoning |
| 45 | Qwen3.6 PlusAlibaba | 50.5 | LMAATBARCLB2/5 | Show details for Qwen3.6 Plus |
| 46 | MiniMax M2.5MiniMax | 50.4 | LMAATBARCLB2/5 | Show details for MiniMax M2.5 |
| 47 | DeepSeek R1DeepSeek | 50.3 | LMAATBARCLB2/5 | Show details for DeepSeek R1 |
| 48 | inklingThinking Machines | 50.3 | LMAATBARCLB✱4/5 | Show details for inkling |
| 49 | GPT-5 MiniOpenAI | 50.2 | LMAATBARCLB2/5 | Show details for GPT-5 Mini |
| 50 | Opus 4.5Anthropic | 49.0 | LMAATBARCLB2/5 | Show details for Opus 4.5 |
| 51 | O3 MiniOpenAI | 48.8 | LMAATBARCLB2/5 | Show details for O3 Mini |
| 52 | Grok 3 MiniSpaceXAI | 48.6 | LMAATBARCLB2/5 | Show details for Grok 3 Mini |
| 53 | MiniMax M3MiniMax | 48.4 | LMAATBARCLB✱3/5 | Show details for MiniMax M3 |
| 54 | Muse GlimmerMeta | 48.3 | LMAATBARCLB2/5 | Show details for Muse Glimmer |
| 55 | GPT-5.4 MiniOpenAI | 47.8 | LMAATBARCLB3/5 | Show details for GPT-5.4 Mini |
| 56 | GPT-5 NanoOpenAI | 47.5 | LMAATBARCLB2/5 | Show details for GPT-5 Nano |
| 57 | GPT-5.4 NanoOpenAI | 47.3 | LMAATBARCLB3/5 | Show details for GPT-5.4 Nano |
| 58 | O1 MiniOpenAI | 47.3 | LMAATBARCLB2/5 | Show details for O1 Mini |
| 59 | Grok 4.3SpaceXAI | 44.8 | LMAATBARCLB2/5 | Show details for Grok 4.3 |
| 60 | Gemini 3.5 Flash LiteGoogle | 43.7 | LMAATBARCLB4/5 | Show details for Gemini 3.5 Flash Lite |
| 61 | Mistral Large 3Mistral | 42.8 | LMAATBARCLB2/5 | Show details for Mistral Large 3 |
| 62 | GPT OSS 120BOpenAI | 42.2 | LMAATBARCLB2/5 | Show details for GPT OSS 120B |
| 63 | Claude 4.5 HaikuAnthropic | 39.5 | LMAATBARCLB2/5 | Show details for Claude 4.5 Haiku |
| — | Muse SparkMetaprovisional | 60.1 | LMAATBARCLB1/5 | Show details for Muse Spark |
| — | GPT-5.2 Chat Latest 20260210OpenAIprovisional | 59.4 | LMAATBARCLB1/5 | Show details for GPT-5.2 Chat Latest 20260210 |
| — | Gemini 3 Deep Think (2/26)Googleprovisional | 59.3 | LMAATBARCLB1/5 | Show details for Gemini 3 Deep Think (2/26) |
| — | Grok 4.20 Beta1SpaceXAIprovisional | 59.3 | LMAATBARCLB1/5 | Show details for Grok 4.20 Beta1 |
| — | GPT-5.5 ProOpenAIprovisional | 59.2 | LMAATBARCLB1/5 | Show details for GPT-5.5 Pro |
| — | Gemini 3 FlashGoogleprovisional | 59.2 | LMAATBARCLB1/5 | Show details for Gemini 3 Flash |
| — | GPT-5.5 InstantOpenAIprovisional | 59.2 | LMAATBARCLB1/5 | Show details for GPT-5.5 Instant |
| — | Qwen3.7 Max PreviewAlibabaprovisional | 59.2 | LMAATBARCLB1/5 | Show details for Qwen3.7 Max Preview |
| — | Claude Opus 4.5 20251101Anthropicprovisional | 59.2 | LMAATBARCLB1/5 | Show details for Claude Opus 4.5 20251101 |
| — | Grok 4.20 Beta 0309 ReasoningSpaceXAIprovisional | 59.1 | LMAATBARCLB1/5 | Show details for Grok 4.20 Beta 0309 Reasoning |
| — | GPT-5.4 ProOpenAIprovisional | 59.1 | LMAATBARCLB1/5 | Show details for GPT-5.4 Pro |
| — | Grok 4.20 Multi Agent Beta 0309SpaceXAIprovisional | 59.0 | LMAATBARCLB1/5 | Show details for Grok 4.20 Multi Agent Beta 0309 |
| — | Ernie 5.1Baiduprovisional | 58.9 | LMAATBARCLB1/5 | Show details for Ernie 5.1 |
| — | GLM-5.1Z.aiprovisional | 58.7 | LMAATBARCLB1/5 | Show details for GLM-5.1 |
| — | Qwen3.5 Max PreviewAlibabaprovisional | 58.7 | LMAATBARCLB1/5 | Show details for Qwen3.5 Max Preview |
| — | Grok 4.1SpaceXAIprovisional | 58.7 | LMAATBARCLB1/5 | Show details for Grok 4.1 |
| — | DeepSeek V4 Pro High 20260813DeepSeekprovisional | 58.4 | LMAATBARCLB1/5 | Show details for DeepSeek V4 Pro High 20260813 |
| — | Qwen3.6 Max PreviewAlibabaprovisional | 58.4 | LMAATBARCLB1/5 | Show details for Qwen3.6 Max Preview |
| — | DeepSeek V4 ProDeepSeekprovisional | 58.2 | LMAATBARCLB1/5 | Show details for DeepSeek V4 Pro |
| — | Dola Seed 2.0 ProByteDanceprovisional | 58.1 | LMAATBARCLB1/5 | Show details for Dola Seed 2.0 Pro |
| — | Claude Sonnet 4.5 20250929Anthropicprovisional | 58.1 | LMAATBARCLB1/5 | Show details for Claude Sonnet 4.5 20250929 |
| — | Qwen3.7 PlusAlibabaprovisional | 58.1 | LMAATBARCLB1/5 | Show details for Qwen3.7 Plus |
| — | DeepSeek V4 Pro High PreviewDeepSeekprovisional | 58.1 | LMAATBARCLB1/5 | Show details for DeepSeek V4 Pro High Preview |
| — | hy3Tencentprovisional | 58.0 | LMAATBARCLB1/5 | Show details for hy3 |
| — | Gemma 4 31BGoogleprovisional | 57.8 | LMAATBARCLB1/5 | Show details for Gemma 4 31B |
| — | Claude 4.7Anthropicprovisional | 57.8 | LMAATBARCLB1/5 | Show details for Claude 4.7 |
| — | Claude Opus 4.1 20250805Anthropicprovisional | 57.7 | LMAATBARCLB1/5 | Show details for Claude Opus 4.1 20250805 |
| — | GPT-5.3 Chat LatestOpenAIprovisional | 57.7 | LMAATBARCLB1/5 | Show details for GPT-5.3 Chat Latest |
| — | Ernie 5.0 Preview 1203Baiduprovisional | 57.7 | LMAATBARCLB1/5 | Show details for Ernie 5.0 Preview 1203 |
| — | Mimo V2 ProXiaomiprovisional | 57.6 | LMAATBARCLB1/5 | Show details for Mimo V2 Pro |
| — | Ernie 5.0 0110Baiduprovisional | 57.5 | LMAATBARCLB1/5 | Show details for Ernie 5.0 0110 |
| — | GPT-4.5 Preview 2025 02 27OpenAIprovisional | 57.4 | LMAATBARCLB1/5 | Show details for GPT-4.5 Preview 2025 02 27 |
| — | ChatGPT-4o Latest 20250326OpenAIprovisional | 57.3 | LMAATBARCLB1/5 | Show details for ChatGPT-4o Latest 20250326 |
| — | GPT-5.2 (Refine.)OpenAIprovisional | 57.3 | LMAATBARCLB1/5 | Show details for GPT-5.2 (Refine.) |
| — | GLM-4.7Z.aiprovisional | 57.2 | LMAATBARCLB1/5 | Show details for GLM-4.7 |
| — | Qwen3.5 397B A17BAlibabaprovisional | 57.2 | LMAATBARCLB1/5 | Show details for Qwen3.5 397B A17B |
| — | DeepSeek V4 Flash High PreviewDeepSeekprovisional | 57.0 | LMAATBARCLB1/5 | Show details for DeepSeek V4 Flash High Preview |
Scores are 0–100 standardized: 50 is average across the models listed, not an absolute grade. Thin evidence is pulled toward 50. A model scored by fewer than 2 recognised boards is listed as provisional without a rank; figures from reports ✱ shape category and domain columns only. A filled badge means that source scored the model; hover for the figure. Open a row for the model card: scores by domain, how to call it, and every source figure. The first 100 of 443 are shown; narrow the filters to see the rest.
§ 2
Recognised leaderboards form the Overall index. Reports ✱ widen the domain columns. A catalog says where a model can be called. All first-party; every figure links to the page it was read from.
Preference
7,999,020 votes across 400 models
Blind head-to-head votes. Measures what people prefer reading, which tracks style as well as correctness.
General intelligence
A composite of ten evaluations Artificial Analysis runs itself, under one harness. Only fully evaluated models are indexed: rows the site stars as estimates (not every component run independently, or based on the lab’s own claims) are left out. Several components (Terminal-Bench among them) are also indexed on their own boards, so the two overlap.
Agentic coding
Top entry dated Sep 3, 2026
Scores a model plus its agent scaffold (Codex, Claude Code, Grok Build). The scaffold is part of the result, not a constant.
Abstract reasoning
Reported per reasoning-effort tier. The same model spans a wide range across tiers; the highest published tier is indexed and recorded beside the score.
General · Reasoning · Code generation · Agentic coding · Mathematics · Data analysis · Language · Instruction following
Questions are refreshed to resist contamination. The overall figure averages the publisher’s own categories, each of which is also indexed on its own.
Report by Institute of Foundation Models (MBZUAI) · 27 benchmarks across 5 categories · marked ✱ wherever it appears
Vendor launch post. Competitor figures are as IFM printed them; the post does not say whether they were re-run or taken from the competitors’ own reports. Effort tiers vary by column.
Catalog · 146 of 465 models matched · ids, context, prices, open weights · never scored
Used as a catalog of where a model can be called and on what terms, not as a source of scores. Prices and availability are the router’s at the date read and move often. Listing here is not an endorsement of the router over the vendor’s own API.
Excluded — SWE-bench VerifiedRead on 2026-09-05: the newest of its 181 entries is dated 2026-02-17. None of the models above appear on it. Including a stale board would add a column of blanks, not a signal. source
§ 3
Five steps, each chosen so that a number here can be traced to a number there, and so that thin or self-reported evidence cannot buy a rank.
Each figure a source publishes is z-scored across the models on it and mapped to 0–100 with 50 as that measure’s average. An Elo of 1504 and a 57.9% resolution rate become comparable, and a narrow-spread board is not drowned out by a wide one.
Where a publisher reports an error bar, the score is discounted in proportion to how wide it is relative to the board’s spread. No error bar means face value — absence is not treated as evidence of a wide one.
A model absent from a source is excluded from that term, not filled in. But thin evidence is pulled toward 50 by a prior worth 25% of the weight available, so one generous board cannot put a barely-tested model at the top. Scored by fewer than 2 recognised boards, a model is listed as provisional and not ranked.
Each recognised board’s headline figure carries a fixed share of the Overall index, stated beside the table with its reason. A board’s category figures shape only the domain they measure, so no board is counted twice.
A launch post or a blog is evidence of a different grade: the publisher chose the benchmarks, the settings and the comparison set. Its figures are indexed for domain and category columns at 10% of a board’s share, never enter the Overall index, and never make a model rankable.
§ 4
An aggregate hides the disagreements that produced it. These are the ones worth knowing before you cite this page.
Sources report several effort tiers per model. The highest published tier is indexed, and the exact label is kept beside each score. Where a source only published a lower tier, that row is weaker evidence than it looks.
Terminal-Bench results are a model plus an agent framework — Codex, Claude Code, Grok Build. The framework is part of the number and is recorded beside each score.
Five sources spell the same model five ways. They are matched by name after tier and vendor words are removed. The rule is tested against hand-checked cases and every merge is logged, but a wrong merge is possible; the model card shows the exact source labels.
A launch post does not always say whether competitor figures were re-run or copied, and its columns can mix effort tiers. They are shown with their mark and their publisher so the reader can weigh them, not laundered into board figures.
Ids, context windows and prices are read from a public catalog on the date shown and move often. Vendor sites are curated pointers. None of it is an endorsement, and none of it touches a score.
Figures were read from each publisher on the date shown. Leaderboards move; check the source before acting on a number.