AI Benchmarks
257 text and coding models, ranked. Context window for 202, input price for 209 and an LMArena Elo for 244 — every figure sourced from the vendor or LMArena, with gaps shown as —.
347
Models
55
Providers
12
Benchmarks
Sep 19, 2026
Data Refreshed
Full Leaderboard
Every text and coding model — sort any column, 10 per page.
Every column is sourced from the vendor or LMArena - sort any column, or search the 257 models below. LMArena Elo shows each row's own retrieval date; the board moves weekly.
| Claude Fable 5 | 1506Sep 13 | $10 | 1M tokens |
| Claude Opus 4.6 (High) | 1505Sep 3 | $5 | 1M tokens |
| Claude Opus 4.7 (High) | 1502Sep 13 | $5 | 1M tokens |
| Muse Spark 1.2 (xHigh) | 1500Sep 14 | $1.25 | 1M tokens |
| Claude Fable 5.1 (Max) | 1498Sep 14 | $10 | 1M tokens |
| Claude Opus 4.6 | 1497Sep 14 | $5 | 1M tokens |
| Claude Opus 4.7 | 1494Sep 14 | $5 | 1M tokens |
| Claude Opus 5 (High) | 1493Sep 14 | $5 | 1M tokens |
| Gemini 3.8 Flash (High) | 1493Sep 14 | $0.75 | 1M tokens |
| Muse Spark 1.1 | 1493Sep 14 | $1.25 | — |
Agent Arena
Agent Arena
Net improvement on agentic tasks, against an average model
LMArena, CC-BY · higher is better · 0% is an average model · top 10 of the 33 rated models carried here
| # | Model | Net |
|---|---|---|
| 1 | Claude Fable 5.1 (Max) | +13.9%±1.95 |
| 2 | GPT-6 Astra (Max) | +12.4%±2.6 |
| 3 | Claude Opus 5 (High) | +11.1%±1.7 |
| 4 | Claude Opus 5 (Max) | +10.8%±1.9 |
| 5 | Claude Opus 4.8 (High) | +7.6%±1.4 |
| 6 | GPT 5.6 Sol (xHigh) | +7.5%±1.4 |
| 7 | Kimi K3 (Max) | +6.5%±0.7 |
| 8 | Claude Sonnet 5 (High) | +5.9%±1.8 |
| 9 | GLM 5.2 (Max) | +4.6%±0.75 |
| 10 | Qwen3.8-Max | +4%±1.05 |
LMArena category ranks
LMArena category ranks
Rank within each category, out of every model LMArena rates
LMArena, CC-BY · lower is better · sorted by Overall
| Model | ||||||||
|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 1 | 1 | 2 | 1 | 1 | 1 | 2 | 2 |
| Claude Opus 4.6 (High) | 2 | 2 | 1 | 3 | 6 | 2 | 1 | 1 |
| Claude Opus 4.7 (High) | 3 | 6 | 4 | 2 | 11 | 5 | 3 | 4 |
| Muse Spark 1.2 (xHigh) | 4 | 27 | 11 | 9 | — | 31 | 18 | 10 |
| Claude Fable 5.1 (Max) | 5 | 8 | 8 | 32 | 4 | 6 | 6 | 5 |
| Claude Opus 4.6 | 6 | 5 | 3 | 5 | 8 | 10 | 4 | 3 |
| Claude Opus 4.7 | 7 | 4 | 5 | 4 | 19 | 8 | 8 | 8 |
| Muse Spark 1.3 (Max) | 8 | 7 | 9 | 8 | 16 | 30 | 15 | 9 |
| Gemini 3.8 Flash (High) | 9 | 31 | 7 | 10 | 9 | 3 | 9 | 6 |
| Claude Opus 5 (High) | 10 | 3 | 6 | 12 | 2 | 11 | 5 | 7 |
| Muse Spark 1.1 | 11 | 28 | 13 | 13 | 24 | 34 | 28 | 39 |
| Gemini 3.7 Flash (High) | 12 | 16 | 18 | 26 | 5 | 4 | 12 | 16 |
| Muse Spark | 13 | 46 | 19 | 19 | 54 | 20 | 38 | 47 |
| Claude Opus 5 (Max) | 14 | 13 | 14 | 18 | 3 | 13 | 7 | 15 |
| Gemini 3.1 Pro | 15 | 22 | 16 | 28 | 22 | 9 | 16 | 13 |
Rankings — Text / LLM
Comparing two models in particular?
The head-to-head tools, category deep dives and written comparisons live on their own page.
Introduction
This page compares 257 current AI models on figures that can be checked: API pricing and context window from each vendor, LMArena human-preference Elo, and six published benchmarks. The table below lists every model individually, ranked by Elo. Nothing on this page is our own scoring.
Evaluation Methodology
The percentages in the charts above are published benchmark results, each from a named source: GPQA Diamond and Mock AIME are run by Epoch AI and released under CC-BY; SimpleBench, Aider Polyglot, LiveBench and the creative-writing benchmark come from their own public leaderboards. Every figure carries the source it was read from.
They are comparable down a column - between models on the same test - and not across columns, because each benchmark measures something different and scores on its own scale. A model with no bar has not been measured on that benchmark rather than scoring zero on it.
For beginners: a benchmark result tells you how a model did on one specific test, which is narrower than "how good is it" and far more checkable. Read several together, notice what each one actually measures, and treat a single number - on this site or any other - as one piece of evidence rather than a verdict.
If You're Optimizing For…
Our picks below, each tied to a fact we can source - not a vendor or third-party claim.
Lowest cost per token
Solar Pro4
$0.03 / $0.12 per million tokens - the lowest first-party vendor API rate on this page
Lowest cost if you shop hosts
Nemotron 3 Super
$0 / $0 at one host - open weights, so the price depends entirely on who serves it
Largest single input
Llama 4 Scout
10M tokens context window, far beyond any other model covered here
Highest LMArena rank
Claude Fable 5
Elo 1506, rank 1 - the highest of the 244 models on this page with a verified leaderboard entry
Self-hosting required
GLM 5.3 (Max), GLM 5.2 (Max), DeepSeek V4 Pro (High), GLM 5, HY3, Gemma 4 31B, Kimi K2.5 (Thinking), DeepSeek V4 Flash High Preview, Gemma 4 26B A4b, GLM 4.7 and 101 more
Top-rated of the 111 models here with published open weights
Smallest hardware footprint
Nemotron 3.5 Lightning
30B parameters with 3B active, documented by NVIDIA as running on a single H100 or A100 80GB
Read Individual Model Guides
Frequently Asked Questions
Which model has the lowest API price?
Solar Pro4, at $0.03 input / $0.12 output per million tokens - the lowest first-party vendor API price among the 257 models on this page. A cheaper headline number exists: Nemotron 3 Super lists $0 / $0, but that is a third-party host's rate, not the model maker's. Open-weight models have no single canonical price - the same weights cost different amounts on different hosts - so we rank them separately rather than against vendors' own API rates. Note DeepSeek's rates also vary by time of day (peak 01:00-04:00 and 06:00-10:00 UTC, half price outside those windows).
Which model has the largest context window?
Llama 4 Scout, at 10M tokens - by far the largest of any model covered here. A large group sits at around 1 million tokens, and Grok 4.5 and Grok 4.6 both offer 500k. Google does not publish a context window figure for Gemini 3.1 Pro, so it is not included in this comparison.
Which model ranks highest on LMArena?
Claude Fable 5, at 1506 (rank 1) - the highest of any model on this page with a leaderboard entry. 244 of the 257 models here have a verified LMArena entry; the rest have no entry we can confirm (including newly-released models), so we leave those blank rather than estimating.
Which models can I actually self-host?
111 models here publish open weights, topped by GLM 5.3 (Max) (MIT), GLM 5.2 (Max) (MIT), DeepSeek V4 Pro (High) (MIT), GLM 5 (MIT), HY3 (Apache 2.0), Gemma 4 31B (Apache 2.0), Kimi K2.5 (Thinking) (Modified MIT), DeepSeek V4 Flash High Preview (MIT), Gemma 4 26B A4b (Apache 2.0), GLM 4.7 (MIT) and 101 more. Everything else here is closed or API-only. Two cases are worth knowing: DeepSeek's V4 models are a paid API with no published open-source licence despite often being described as open-weight, and GLM-5.3 withheld its weights at launch even though GLM-5.2's are published under MIT.
Which model is preview vs. stable?
Gemini 3.1 Pro (preview) and Qwen3.8-Flash-Next (open-weight preview of the qwen4 architecture) are not documented as stable releases. The other 255 models covered here are documented as regular releases by their vendors.
Where do the benchmark percentages come from?
Each column is one published benchmark, as a percentage: GPQA Diamond and Mock AIME from Epoch AI, SimpleBench, Aider Polyglot, LiveBench and a creative-writing benchmark from their own leaderboards. They are comparable down a column, between models. They are not comparable across columns - 62 on GPQA and 62 on Aider are two different tests. Nothing in these charts is our own scoring.