Skip to content

AI Benchmarks

257 text and coding models, ranked. Context window for 202, input price for 209 and an LMArena Elo for 244 — every figure sourced from the vendor or LMArena, with gaps shown as —.

347

Models

55

Providers

12

Benchmarks

Sep 19, 2026

Data Refreshed

Full Leaderboard

Every text and coding model — sort any column, 10 per page.

Every column is sourced from the vendor or LMArena - sort any column, or search the 257 models below. LMArena Elo shows each row's own retrieval date; the board moves weekly.

257 AI models. LMArena Elo, input price per million tokens and context window. Sortable by any column.
Claude Fable 51506Sep 13$101M tokens
Claude Opus 4.6 (High)1505Sep 3$51M tokens
Claude Opus 4.7 (High)1502Sep 13$51M tokens
Muse Spark 1.2 (xHigh)1500Sep 14$1.251M tokens
Claude Fable 5.1 (Max)1498Sep 14$101M tokens
Claude Opus 4.61497Sep 14$51M tokens
Claude Opus 4.71494Sep 14$51M tokens
Claude Opus 5 (High)1493Sep 14$51M tokens
Gemini 3.8 Flash (High)1493Sep 14$0.751M tokens
Muse Spark 1.11493Sep 14$1.25
1-10 of 257
Page 1 of 26

Agent Arena

Agent Arena

Net improvement on agentic tasks, against an average model

LMArena, CC-BY · higher is better · 0% is an average model · top 10 of the 33 rated models carried here

33 AI models by LMArena Agent Arena net improvement, in percentage points against an average model. Positive is above average, negative below.
#ModelNet
1Claude Fable 5.1 (Max)+13.9%±1.95
2GPT-6 Astra (Max)+12.4%±2.6
3Claude Opus 5 (High)+11.1%±1.7
4Claude Opus 5 (Max)+10.8%±1.9
5Claude Opus 4.8 (High)+7.6%±1.4
6GPT 5.6 Sol (xHigh)+7.5%±1.4
7Kimi K3 (Max)+6.5%±0.7
8Claude Sonnet 5 (High)+5.9%±1.8
9GLM 5.2 (Max)+4.6%±0.75
10Qwen3.8-Max+4%±1.05

LMArena category ranks

LMArena category ranks

Rank within each category, out of every model LMArena rates

LMArena, CC-BY · lower is better · sorted by Overall

215 AI models ranked by LMArena in 8 categories. First, second and third place are shaded; the rank is printed in every cell. An em dash means LMArena does not rank that model in that category.
Model
Claude Fable 511211122
Claude Opus 4.6 (High)22136211
Claude Opus 4.7 (High)364211534
Muse Spark 1.2 (xHigh)427119311810
Claude Fable 5.1 (Max)588324665
Claude Opus 4.6653581043
Claude Opus 4.7745419888
Muse Spark 1.3 (Max)87981630159
Gemini 3.8 Flash (High)9317109396
Claude Opus 5 (High)10361221157
Muse Spark 1.11128131324342839
Gemini 3.7 Flash (High)12161826541216
Muse Spark1346191954203847
Claude Opus 5 (Max)14131418313715
Gemini 3.1 Pro152216282291613
115 of 215
Showing

Rankings — Text / LLM

Comparing two models in particular?

The head-to-head tools, category deep dives and written comparisons live on their own page.

Go to comparisons

Introduction

This page compares 257 current AI models on figures that can be checked: API pricing and context window from each vendor, LMArena human-preference Elo, and six published benchmarks. The table below lists every model individually, ranked by Elo. Nothing on this page is our own scoring.

Evaluation Methodology

The percentages in the charts above are published benchmark results, each from a named source: GPQA Diamond and Mock AIME are run by Epoch AI and released under CC-BY; SimpleBench, Aider Polyglot, LiveBench and the creative-writing benchmark come from their own public leaderboards. Every figure carries the source it was read from.

They are comparable down a column - between models on the same test - and not across columns, because each benchmark measures something different and scores on its own scale. A model with no bar has not been measured on that benchmark rather than scoring zero on it.

For beginners: a benchmark result tells you how a model did on one specific test, which is narrower than "how good is it" and far more checkable. Read several together, notice what each one actually measures, and treat a single number - on this site or any other - as one piece of evidence rather than a verdict.

If You're Optimizing For…

Our picks below, each tied to a fact we can source - not a vendor or third-party claim.

Lowest cost per token

Solar Pro4

$0.03 / $0.12 per million tokens - the lowest first-party vendor API rate on this page

Lowest cost if you shop hosts

Nemotron 3 Super

$0 / $0 at one host - open weights, so the price depends entirely on who serves it

Largest single input

Llama 4 Scout

10M tokens context window, far beyond any other model covered here

Highest LMArena rank

Claude Fable 5

Elo 1506, rank 1 - the highest of the 244 models on this page with a verified leaderboard entry

Self-hosting required

GLM 5.3 (Max), GLM 5.2 (Max), DeepSeek V4 Pro (High), GLM 5, HY3, Gemma 4 31B, Kimi K2.5 (Thinking), DeepSeek V4 Flash High Preview, Gemma 4 26B A4b, GLM 4.7 and 101 more

Top-rated of the 111 models here with published open weights

Smallest hardware footprint

Nemotron 3.5 Lightning

30B parameters with 3B active, documented by NVIDIA as running on a single H100 or A100 80GB

Read Individual Model Guides

View all 257 models

Frequently Asked Questions

Which model has the lowest API price?

Solar Pro4, at $0.03 input / $0.12 output per million tokens - the lowest first-party vendor API price among the 257 models on this page. A cheaper headline number exists: Nemotron 3 Super lists $0 / $0, but that is a third-party host's rate, not the model maker's. Open-weight models have no single canonical price - the same weights cost different amounts on different hosts - so we rank them separately rather than against vendors' own API rates. Note DeepSeek's rates also vary by time of day (peak 01:00-04:00 and 06:00-10:00 UTC, half price outside those windows).

Which model has the largest context window?

Llama 4 Scout, at 10M tokens - by far the largest of any model covered here. A large group sits at around 1 million tokens, and Grok 4.5 and Grok 4.6 both offer 500k. Google does not publish a context window figure for Gemini 3.1 Pro, so it is not included in this comparison.

Which model ranks highest on LMArena?

Claude Fable 5, at 1506 (rank 1) - the highest of any model on this page with a leaderboard entry. 244 of the 257 models here have a verified LMArena entry; the rest have no entry we can confirm (including newly-released models), so we leave those blank rather than estimating.

Which models can I actually self-host?

111 models here publish open weights, topped by GLM 5.3 (Max) (MIT), GLM 5.2 (Max) (MIT), DeepSeek V4 Pro (High) (MIT), GLM 5 (MIT), HY3 (Apache 2.0), Gemma 4 31B (Apache 2.0), Kimi K2.5 (Thinking) (Modified MIT), DeepSeek V4 Flash High Preview (MIT), Gemma 4 26B A4b (Apache 2.0), GLM 4.7 (MIT) and 101 more. Everything else here is closed or API-only. Two cases are worth knowing: DeepSeek's V4 models are a paid API with no published open-source licence despite often being described as open-weight, and GLM-5.3 withheld its weights at launch even though GLM-5.2's are published under MIT.

Which model is preview vs. stable?

Gemini 3.1 Pro (preview) and Qwen3.8-Flash-Next (open-weight preview of the qwen4 architecture) are not documented as stable releases. The other 255 models covered here are documented as regular releases by their vendors.

Where do the benchmark percentages come from?

Each column is one published benchmark, as a percentage: GPQA Diamond and Mock AIME from Epoch AI, SimpleBench, Aider Polyglot, LiveBench and a creative-writing benchmark from their own leaderboards. They are comparable down a column, between models. They are not comparable across columns - 62 on GPQA and 62 on Aider are two different tests. Nothing in these charts is our own scoring.