ModelLab
ModelsCompare
Updated Jun 2026

Every frontier AI model,
one clear leaderboard.

Compare models from OpenAI, Anthropic, Google, xAI, Meta, DeepSeek, Qwen and Mistral. Filter by lab, sort by the benchmark that matters, and line them up side-by-side.

Explore modelsCompare side-by-side

17

Models tracked

Claude Opus 4.8

Current Arena leader

Claude Opus 4.8

Newest release

5

Open-weights models

Leaderboard

The model lineup

Sort by any benchmark, filter by lab or tier. Scores are the best reported figures — hover a column header for what it measures.

▾
Lab
Tier
17 models
#ModelArenaSWEAIMEGPQAMMLUAiderOSContextPrice (in/out)
A
Claude Opus 4.8

Anthropic · Jun 3, 2026

152181.295.890.487.988.544.1500K$15 / $75
2OA
GPT-5.5

OpenAI · May 14, 2026

151274.996.489.186.784.242.0400K$2.50 / $10
3G
Gemini 3.1 Pro

Google · May 20, 2026

150572.097.191.286.480.340.62M$2.50 / $10
4X
Grok 4.3

xAI · May 28, 2026

149875.895.087.385.083.036.5256K$3.00 / $15
5A
Claude Sonnet 4.6

Anthropic · Mar 18, 2026

148673.492.084.584.082.138.9500K$3.00 / $15
6∞
Llama 4 BehemothOpen

Meta · Apr 29, 2026

146966.092.583.183.577.428.51MOpen
7DS
DeepSeek V4Open

DeepSeek · Apr 12, 2026

146164.593.282.482.074.025.0128KOpen
8OA
GPT-5.1

OpenAI · Oct 20, 2025

145868.594.288.085.179.638.4400K$1.50 / $6.00
9G
Gemini 3 Flash

Google · Apr 8, 2026

145160.190.482.082.073.530.21M$0.30 / $2.50
10Q
Qwen3-Max

Qwen · Mar 24, 2026

144458.989.080.580.270.521.5256K$0.50 / $2.00
11M
Mistral Large 3

Mistral · Feb 25, 2026

142952.682.078.078.567.216.5128K$2.00 / $6.00
12X
Grok 4 Fast

xAI · Mar 2, 2026

142249.786.076.878.566.218.0131K$0.20 / $1.50
13∞
Llama 4 MaverickOpen

Meta · Jan 15, 2026

141754.384.077.979.069.820.01MOpen
14Q
Qwen3-235BOpen

Qwen · Feb 7, 2026

140848.281.075.076.863.014.0131KOpen
15OA
GPT-5.1 mini

OpenAI · Nov 4, 2025

140251.288.578.378.968.422.0200K$0.40 / $1.60
16A
Claude Haiku 4.6

Anthropic · Feb 11, 2026

139544.880.171.275.461.015.5250K$0.80 / $4.00
17M
Codestral 2Open

Mistral · Jan 30, 2026

137247.0—65.070.076.5—256KOpen
A

Claude Opus 4.8

Top of the LLM Stats overall board; leading on deep coding.

AnthropicFrontier

Arena

1521

Arena

1521

SWE

81.2

AIME

95.8

GPQA

90.4

MMLU

87.9

Aider

88.5

OS

44.1

500K ctx$15/$75
Details
OA

GPT-5.5

OpenAI's flagship all-rounder with strong computer use.

OpenAIFrontier

Arena

1512

Arena

1512

SWE

74.9

AIME

96.4

GPQA

89.1

MMLU

86.7

Aider

84.2

OS

42.0

400K ctx$2.50/$10
Details
G

Gemini 3.1 Pro

Leads on reasoning with a native 2M-token context.

GoogleFrontier

Arena

1505

Arena

1505

SWE

72.0

AIME

97.1

GPQA

91.2

MMLU

86.4

Aider

80.3

OS

40.6

2M ctx$2.50/$10
Details
X

Grok 4.3

Real-time X integration, top-tier coding and reasoning.

xAIFrontier

Arena

1498

Arena

1498

SWE

75.8

AIME

95.0

GPQA

87.3

MMLU

85.0

Aider

83.0

OS

36.5

256K ctx$3.00/$15
Details
A

Claude Sonnet 4.6

The model most developers actually ship on.

AnthropicFrontier

Arena

1486

Arena

1486

SWE

73.4

AIME

92.0

GPQA

84.5

MMLU

84.0

Aider

82.1

OS

38.9

500K ctx$3.00/$15
Details
∞

Llama 4 Behemoth

Meta's largest open-weights model, near-frontier quality.

MetaFrontierOpen weights

Arena

1469

Arena

1469

SWE

66.0

AIME

92.5

GPQA

83.1

MMLU

83.5

Aider

77.4

OS

28.5

1M ctxOpen
Details
DS

DeepSeek V4

Frontier-class reasoning at a fraction of the price.

DeepSeekFrontierOpen weights

Arena

1461

Arena

1461

SWE

64.5

AIME

93.2

GPQA

82.4

MMLU

82.0

Aider

74.0

OS

25.0

128K ctxOpen
Details
OA

GPT-5.1

Previous flagship — now value-priced, still excellent.

OpenAIFrontier

Arena

1458

Arena

1458

SWE

68.5

AIME

94.2

GPQA

88.0

MMLU

85.1

Aider

79.6

OS

38.4

400K ctx$1.50/$6.00
Details
G

Gemini 3 Flash

Best price-to-performance in the lineup, huge context.

GoogleMid-tier

Arena

1451

Arena

1451

SWE

60.1

AIME

90.4

GPQA

82.0

MMLU

82.0

Aider

73.5

OS

30.2

1M ctx$0.30/$2.50
Details
Q

Qwen3-Max

Top open-source lineage; best-in-class multilingual.

QwenFrontier

Arena

1444

Arena

1444

SWE

58.9

AIME

89.0

GPQA

80.5

MMLU

80.2

Aider

70.5

OS

21.5

256K ctx$0.50/$2.00
Details
M

Mistral Large 3

Europe's flagship — strong, efficient, GDPR-friendly.

MistralFrontier

Arena

1429

Arena

1429

SWE

52.6

AIME

82.0

GPQA

78.0

MMLU

78.5

Aider

67.2

OS

16.5

128K ctx$2.00/$6.00
Details
X

Grok 4 Fast

Reasoning model tuned for low latency.

xAISmall / fast

Arena

1422

Arena

1422

SWE

49.7

AIME

86.0

GPQA

76.8

MMLU

78.5

Aider

66.2

OS

18.0

131K ctx$0.20/$1.50
Details
∞

Llama 4 Maverick

MoE workhorse — efficient and openly deployable.

MetaMid-tierOpen weights

Arena

1417

Arena

1417

SWE

54.3

AIME

84.0

GPQA

77.9

MMLU

79.0

Aider

69.8

OS

20.0

1M ctxOpen
Details
Q

Qwen3-235B

Open-weights MoE with a thriving community ecosystem.

QwenOpen weightsOpen weights

Arena

1408

Arena

1408

SWE

48.2

AIME

81.0

GPQA

75.0

MMLU

76.8

Aider

63.0

OS

14.0

131K ctxOpen
Details
OA

GPT-5.1 mini

Fast, cheap, capable — the everyday workhorse.

OpenAISmall / fast

Arena

1402

Arena

1402

SWE

51.2

AIME

88.5

GPQA

78.3

MMLU

78.9

Aider

68.4

OS

22.0

200K ctx$0.40/$1.60
Details
A

Claude Haiku 4.6

Sub-second responses at near-frontier quality.

AnthropicSmall / fast

Arena

1395

Arena

1395

SWE

44.8

AIME

80.1

GPQA

71.2

MMLU

75.4

Aider

61.0

OS

15.5

250K ctx$0.80/$4.00
Details
M

Codestral 2

Code-specialist tuned for fast IDE autocomplete.

MistralMid-tierOpen weights

Arena

1372

Arena

1372

SWE

47.0

AIME

—

GPQA

65.0

MMLU

70.0

Aider

76.5

OS

—

256K ctxOpen
Details

Methodology

What the benchmarks measure

Each column maps to a specific, publicly reported eval. Click through to a model for the full breakdown.

Chatbot Arena

Elo

Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.

SWE-bench Verified

0–100

Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.

AIME 2025

0–100

American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.

GPQA Diamond

0–100

Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.

MMLU-Pro

0–100

Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.

Aider Polyglot

0–100

Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).

OSWorld

0–100

Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.

ModelLab

An independent dashboard for comparing frontier AI models. Benchmark scores are drawn from public model cards, papers, and leaderboards. Always verify against the original source before making decisions.

Navigate

  • All models
  • Compare
  • About the data

Sources

  • LMArena ↗
  • SWE-bench ↗
  • LiveBench ↗

Data snapshot last checked June 17, 2026.

Not affiliated with any AI lab. Scores are approximations.