ModelBench
June 2026
Data from LMSYS Arena • Artificial Analysis • Providers
Cost Calculator
LIVE • 13 frontier & strong models

Latest AI models.
Clear benchmarks.

Filter, sort, and compare the best models from every major lab. Find the right model for coding, reasoning, or cost.

Claude Fable 5 leads Arena at 1510
Gemini 3.5 Flash: fastest high-performer
GLM-5.2 & DeepSeek best open value
QUICK START
LABS
MIN ARENA ELO
1440
MAX INPUT PRICE
$12
Showing 13 of 13 models
Sort
Anthropic
Claude Fable 5
1510
ELO
MMLU-Pro91.5
GPQA95
SWE-Bench80.3
AIME99.5
$10 / $50/M
1M ctx•42 t/s
CodingReasoningTop TierThinking
Anthropic
Claude Opus 4.8 Thinking
1506
ELO
MMLU-Pro90.1
GPQA93.8
SWE-Bench78.5
AIME99.2
$5 / $25/M
1M ctx•48 t/s
CodingReasoningTop TierThinking
OpenAI
GPT-5.5 (high)
1506
ELO
MMLU-Pro89.6
GPQA94.4
SWE-Bench76.9
AIME99.8
$5 / $30/M
1M ctx•51 t/s
CodingReasoningTop TierThinking
Google
Gemini 3.1 Pro
1505
ELO
MMLU-Pro91
GPQA94.3
SWE-Bench80.6
AIME100
$2 / $12/M
1M ctx•68 t/s
CodingReasoningTop TierThinking
Anthropic
Claude Opus 4.8
1504
ELO
MMLU-Pro90
GPQA92.4
SWE-Bench77.1
AIME98.8
$5 / $25/M
1M ctx•55 t/s
CodingReasoningTop TierLong Context
Google
Gemini 3.5 Flash
1504
ELO
MMLU-Pro91
GPQA89.2
SWE-Bench71.4
AIME96.4
$0.35 / $1.5/M
1M ctx•142 t/s
Top TierLong ContextValueFast
xAI
Grok-4.20
1496
ELO
MMLU-Pro89.6
GPQA91.2
SWE-Bench75
AIME98
$2.5 / $7.5/M
2M ctx•88 t/s
CodingTop TierThinkingLong Context
Zhipu AI
GLM-5.2
1488
ELO
MMLU-Pro87.5
GPQA89.5
SWE-Bench72.8
AIME97.2
$0.8 / $2.2/M
1M ctx•95 t/s
ThinkingLong ContextValueFastOPEN
Alibaba
Qwen3.7-Max
1486
ELO
MMLU-Pro89.6
GPQA90.8
SWE-Bench71.9
AIME98.5
$1.5 / $6/M
1M ctx•62 t/s
ThinkingLong ContextValueMultimodal
Meta
Muse Spark
1485
ELO
MMLU-Pro87.3
GPQA88.5
SWE-Bench69.1
AIME95
$0.9 / $3.5/M
512K ctx•75 t/s
ValueMultimodal
DeepSeek
DeepSeek-V4 Pro
1467
ELO
MMLU-Pro87.5
GPQA88.8
SWE-Bench73.5
AIME95
$0.55 / $2.1/M
128K ctx•102 t/s
ThinkingValueFastOPEN
OpenAI
GPT-5.4
1465
ELO
MMLU-Pro88.4
GPQA92
SWE-Bench74.8
AIME97.8
$2.5 / $15/M
1M ctx•57 t/s
ReasoningThinkingLong ContextMultimodal
Anthropic
Claude Sonnet 4.6
1460
ELO
MMLU-Pro87.3
GPQA84.1
SWE-Bench79.6
AIME95.5
$3 / $15/M
1M ctx•78 t/s
CodingThinkingLong ContextMultimodal
Monthly Cost Estimator
Assumes no caching
10M
3M
ESTIMATED MONTHLY COST
Claude $250
Claude $125
GPT-5.5 $140
Gemini $56
Claude $125
Gemini $8
Use the Compare view for full side-by-side cost + performance.
Data is approximate and aggregated from public leaderboards (LMSYS / lmarena, Artificial Analysis) and provider announcements as of June 18, 2026. Always verify with official sources before production use.
Built for clarity • Mobile-first design • Compare up to 5 models