BM
BENCHMARKLIVEAug 2026 • 20 models • 9 benchmarks
Independent frontier leaderboard — filter, sort, compare.
Methodology
Updated hourly from official tech reports & LMSYS ArenaPrices per 1M tokens • Context = max supported • Latency p50

Which model actually wins?

Side-by-side benchmarks for GPT, Claude, Gemini, Llama, Mistral, Grok & more. Start from your use case, not the hype — filter by modality, price, and the benchmarks that matter to you.

Reasoning-heavy? Sort by AIME / GPQA
Shipping product? Sort by SWE-Bench + Price
#1 FRONTIER
A
Anthropic
Claude 4 Opus
Claude 4 • 2025-05
1382Arena ELO
SWE
79.4%
GPQA
74.1%
MATH
95.3%
#2
G
Google
Gemini 2.5 Pro
Gemini 2.5 • 2025-03
1365Arena ELO
SWE
63.8%
GPQA
73.5%
MATH
94.2%
#3
O
OpenAI
o3
o-series • 2025-04
1362Arena ELO
SWE
69.1%
GPQA
75.2%
MATH
96.1%
⌕
20 models
LAB
MODALITY
ACCESS
BENCHMARK FOCUS
A
Claude 4 Opus
Anthropic • Claude 4 • 2025-05
textvisioncodereasoning200K • —
ARENA ELO
1382
Chatbot Arena
PRICE / 1M
$15 / $75
in / out
SWE-Bench Verified
79.4%
focus
MMLU
91.1
MMLU-Pro
80.2
GPQA Diamond
74.1
MATH
95.3
AIME 2024
76.6
Complex tasksResearch
G
Gemini 2.5 Pro
Google • Gemini 2.5 • 2025-03
textvisionaudioreasoning1M • —
ARENA ELO
1365
Chatbot Arena
PRICE / 1M
$1.25 / $10
in / out
SWE-Bench Verified
63.8%
focus
MMLU
92
MMLU-Pro
78.9
GPQA Diamond
73.5
MATH
94.2
AIME 2024
70
Video analysisLong docs
O
o3
OpenAI • o-series • 2025-04
textvisionreasoningcode200K • —
ARENA ELO
1362
Chatbot Arena
PRICE / 1M
$10 / $40
in / out
SWE-Bench Verified
69.1%
focus
MMLU
91.7
MMLU-Pro
79.2
GPQA Diamond
75.2
MATH
96.1
AIME 2024
83.3
Hard reasoningResearch
A
Claude 4 Sonnet
Anthropic • Claude 4 • 2025-05
textvisioncodereasoning200K • —
ARENA ELO
1355
Chatbot Arena
PRICE / 1M
$3 / $15
in / out
SWE-Bench Verified
72.7%
focus
MMLU
90.4
MMLU-Pro
77.1
GPQA Diamond
70.3
MATH
92.4
AIME 2024
51.6
Software engineeringAgents
DS
DeepSeek-R1
DeepSeek • DeepSeek R1 • 2025-01
textcodereasoningOPEN128K • 671B (37B active)
ARENA ELO
1348
Chatbot Arena
PRICE / 1M
$0.55 / $2.19
in / out
SWE-Bench Verified
49.2%
focus
MMLU
90.8
MMLU-Pro
77.4
GPQA Diamond
71.5
MATH
97.3
AIME 2024
79.8
Open reasoning
O
GPT-4.5
OpenAI • GPT-4.5 • 2025-02
textvision128K • —
ARENA ELO
1340
Chatbot Arena
PRICE / 1M
$75 / $150
in / out
SWE-Bench Verified
38%
focus
MMLU
90.1
MMLU-Pro
75.4
GPQA Diamond
62.1
MATH
82.3
AIME 2024
36.6
WritingBrainstorming
O
o1
OpenAI • o-series • 2024-09
textvisionreasoning200K • —
ARENA ELO
1335
Chatbot Arena
PRICE / 1M
$15 / $60
in / out
SWE-Bench Verified
48.9%
focus
MMLU
92.3
MMLU-Pro
78.4
GPQA Diamond
78
MATH
94.8
AIME 2024
74
STEMAnalysis
𝕏
Grok 3
xAI • Grok • 2025-02
textvisionreasoning128K • —
ARENA ELO
1332
Chatbot Arena
PRICE / 1M
$3 / $15
in / out
SWE-Bench Verified
47.6%
focus
MMLU
89.2
MMLU-Pro
74.3
GPQA Diamond
66.2
MATH
89.1
AIME 2024
58.3
Social, search
Q
Qwen3 235B
Alibaba • Qwen3 • 2025-04
textcodereasoningOPEN128K • 235B (22B active)
ARENA ELO
1328
Chatbot Arena
PRICE / 1M
$0.22 / $0.88
in / out
SWE-Bench Verified
45.7%
focus
MMLU
89.1
MMLU-Pro
75.2
GPQA Diamond
68.4
MATH
92.8
AIME 2024
68.4
MultilingualValue
M
Llama 4 Maverick
Meta • Llama 4 • 2025-04
textvisioncodeOPEN1M • 400B (17B active)
ARENA ELO
1322
Chatbot Arena
PRICE / 1M
$0.27 / $0.85
in / out
SWE-Bench Verified
46.3%
focus
MMLU
88.4
MMLU-Pro
74.6
GPQA Diamond
60.1
MATH
81
AIME 2024
24
Self-hostFine-tune
DS
DeepSeek-V3
DeepSeek • DeepSeek V3 • 2024-12
textcodeOPEN128K • 671B (37B active)
ARENA ELO
1318
Chatbot Arena
PRICE / 1M
$0.27 / $1.1
in / out
SWE-Bench Verified
42%
focus
MMLU
88.5
MMLU-Pro
75.9
GPQA Diamond
59.1
MATH
90.2
AIME 2024
39.2
Cost-efficient coding
O
GPT-4o
OpenAI • GPT-4o • 2024-05
textvisionaudio128K • —
ARENA ELO
1314
Chatbot Arena
PRICE / 1M
$2.5 / $10
in / out
SWE-Bench Verified
33.2%
focus
MMLU
88.7
MMLU-Pro
72.6
GPQA Diamond
53.6
MATH
76.6
AIME 2024
13.3
General chatMultimodal apps
G
Gemini 2.0 Flash
Google • Gemini 2.0 • 2024-12
textvisionaudio1M • —
ARENA ELO
1283
Chatbot Arena
PRICE / 1M
$0.1 / $0.4
in / out
SWE-Bench Verified
42.1%
focus
MMLU
87.2
MMLU-Pro
69.8
GPQA Diamond
58.2
MATH
78.6
AIME 2024
28.3
Real-time appsScale
A
Claude 3.5 Sonnet
Anthropic • Claude 3.5 • 2024-06
textvisioncode200K • —
ARENA ELO
1272
Chatbot Arena
PRICE / 1M
$3 / $15
in / out
SWE-Bench Verified
49%
focus
MMLU
88.7
MMLU-Pro
73
GPQA Diamond
59.4
MATH
71.1
AIME 2024
16
General purpose
G
Gemini 1.5 Pro
Google • Gemini 1.5 • 2024-02
textvisionaudio2M • —
ARENA ELO
1248
Chatbot Arena
PRICE / 1M
$1.25 / $10
in / out
SWE-Bench Verified
29%
focus
MMLU
85.9
MMLU-Pro
68.1
GPQA Diamond
46.2
MATH
58.5
AIME 2024
8
Long docs
Mi
Mistral Large 2
Mistral • Mistral Large • 2024-07
textcodeOPEN128K • 123B
ARENA ELO
1216
Chatbot Arena
PRICE / 1M
$2 / $6
in / out
SWE-Bench Verified
30.1%
focus
MMLU
84.5
MMLU-Pro
66.8
GPQA Diamond
42.1
MATH
68.2
AIME 2024
10
Enterprise EU
M
Llama 3.3 70B
Meta • Llama 3 • 2024-12
textcodeOPEN128K • 70B
ARENA ELO
1212
Chatbot Arena
PRICE / 1M
$0.12 / $0.12
in / out
SWE-Bench Verified
33%
focus
MMLU
86
MMLU-Pro
69.1
GPQA Diamond
45
MATH
67
AIME 2024
6.6
StartupsEdge
G
Gemma 3 27B
Google • Gemma • 2025-03
textvisionOPEN32K • 27B
ARENA ELO
1188
Chatbot Arena
PRICE / 1M
$0.1 / $0.2
in / out
SWE-Bench Verified
22.1%
focus
MMLU
82.1
MMLU-Pro
63.4
GPQA Diamond
38.2
MATH
56.1
AIME 2024
6
On-deviceEdge
J
Jamba 1.6 Large
AI21 • Jamba • 2025-03
textcode256K • 398B (52B active)
ARENA ELO
1184
Chatbot Arena
PRICE / 1M
$2 / $8
in / out
SWE-Bench Verified
29.4%
focus
MMLU
84.2
MMLU-Pro
66
GPQA Diamond
43.5
MATH
62.1
AIME 2024
8.3
Long context
C
Command R+
Cohere • Command • 2024-04
text128K • 104B
ARENA ELO
1158
Chatbot Arena
PRICE / 1M
$3 / $15
in / out
SWE-Bench Verified
17.3%
focus
MMLU
79.7
MMLU-Pro
62.1
GPQA Diamond
40.2
MATH
50
AIME 2024
3.3
Enterprise RAG

How we score & what to watch

Benchmarks

MMLU / MMLU-Pro for breadth, GPQA Diamond for grad-level science, MATH + AIME for elite math, HumanEval + SWE-Bench Verified for real code, MMMU for vision, IFEval for instruction following.

Pick by job

Reasoning & research: sort by AIME + GPQA. Shipping features: SWE-Bench + latency + price. Long docs/video: filter 200K+ and Gemini/GPT-4o class. Self-host: open-weights + price.

Caveats

Scores are noisy, prompts differ by lab, and contamination happens. Pair benchmarks with Arena ELO (human preference) and your own evals. Prices and context windows change frequently — verify before prod.

Data: official model cards + LMSYS Arena (Aug 20, 2026)Not affiliated with any lab — branding used for identification
© 2026 BENCHMARK — Built mobile-first. Data Aug 2026. Not affiliated with model providers.PrivacyChangelogAPI