Model GaugeAI benchmarks, usable
ExploreCompareMethod

Snapshot 2026-07-15

Pick the right model for the job

Filter by lab and use case, sort by Arena, Artificial Analysis, SWE-bench, or price, then compare up to four models side by side — including Cursor Grok 4.5, who built this site.

Snapshot dated 2026-07-15. Scores change weekly; treat this as a decision aid, not a live API.

19 models · preset Frontier overall

AN
AnthropicSpotlight

Claude Fable 5

Hard coding agents, long computer-use sessions, and high-stakes knowledge work.

Details
Arena
1525
AA Index
60
SWE-bench
95%
GPQA
93%
Context 1MPrice $10 / $50 /M in·outSpeed 58 t/sCoding · Agents · Writing
OP
OpenAISpotlight

GPT-5.6 Sol

General frontier reasoning, mixed multimodal apps, and OpenAI ecosystem tooling.

Details
Arena
1514
AA Index
59
SWE-bench
—
GPQA
—
Context 1MPrice $5 / $30 /M in·outSpeed 57 t/sCoding · Agents · Multimodal
QW
Qwen

Qwen 3.7 Max

Multilingual agents, Chinese-market products, and long agentic sessions.

Details
Arena
1488
AA Index
57
SWE-bench
—
GPQA
—
Context 1MPrice $1.2 / $6 /M in·outSpeed 75 t/sAgents · Coding · Writing
AN
AnthropicSpotlight

Claude Opus 4.8

Coding agents, computer use, and careful long-form reasoning in production.

Details
Arena
1512
AA Index
56
SWE-bench
89%
GPQA
94%
Context 1MPrice $5 / $25 /M in·outSpeed 72 t/sCoding · Agents · Writing
OP
OpenAI

GPT-5.5 Pro

Hard reasoning tasks, research assistants, and multimodal product surfaces.

Details
Arena
1510
AA Index
55
SWE-bench
80%
GPQA
92%
Context 1MPrice $30 / $180 /M in·outSpeed 68 t/sScience · Agents · Multimodal
SP
SpaceXAISpotlight

Cursor Grok 4.5

Coding agents in Cursor / Grok Build, long-running engineering tasks, and near-frontier intelligence at a fraction of Opus/GPT Pro spend.

Details
Arena
—
AA Index
54
SWE-bench
—
GPQA
—
Context 500KPrice $2 / $6 /M in·outCoding · Agents · Writing
GO
GoogleSpotlight

Gemini 3.1 Pro

Scientific reasoning, huge document corpora, and Google Cloud / Workspace stacks.

Details
Arena
1492
AA Index
54
SWE-bench
81%
GPQA
94%
Context 2MPrice $2.5 / $15 /M in·outSpeed 85 t/sScience · Long context · Multimodal
AN
Anthropic

Claude Opus 4.7

Complex coding agents and careful editorial writing.

Details
Arena
1505
AA Index
52
SWE-bench
86%
GPQA
92%
Context 1MPrice $5 / $25 /M in·outSpeed 70 t/sCoding · Agents · Writing
OP
OpenAI

GPT-5.5

Default production chat, tools, and mixed workloads on OpenAI.

Details
Arena
1506
AA Index
51
SWE-bench
78%
GPQA
90%
Context 1MPrice $5 / $30 /M in·outSpeed 70 t/sWriting · Multimodal · Agents
XA
xAI

Grok 4.3

Current-events chat, witty assistants, and X-integrated products.

Details
Arena
1496
AA Index
50
SWE-bench
—
GPQA
—
Context 2MPrice $3 / $15 /M in·outSpeed 110 t/sWriting · Speed · Long context
AN
AnthropicSpotlight

Claude Sonnet 4

Day-to-day coding copilots, customer agents, and high-volume Claude deployments.

Details
Arena
1518
AA Index
48
SWE-bench
82%
GPQA
88%
Context 1MPrice $3 / $15 /M in·outSpeed 95 t/sCoding · Agents · Writing
OP
OpenAI

GPT-5.4

Teams already on GPT-5.x infra who need stability over bleeding edge.

Details
Arena
1463
AA Index
46
SWE-bench
80%
GPQA
92%
Context 1.1MPrice $2.5 / $15 /M in·outSpeed 80 t/sCoding · Agents · Multimodal
GO
Google

Gemini 3.0

Balanced Google Cloud apps needing solid quality at lower spend.

Details
Arena
1505
AA Index
45
SWE-bench
74%
GPQA
88%
Context 1MPrice $1.3 / $5 /M in·outSpeed 100 t/sBest value · Multimodal · Speed
DE
DeepSeekOpenSpotlight

DeepSeek V4 Pro

Self-hosting, cost-sensitive production, and research fine-tunes.

Details
Arena
1462
AA Index
44
SWE-bench
76%
GPQA
88%
Context 164KPrice $0.40 / $1.2 /M in·outSpeed 90 t/sBest value · Coding · Science
MI
MiniMaxOpen

MiniMax M2.5

Open coding agents and competitive self-hosted SWE workflows.

Details
Arena
—
AA Index
44
SWE-bench
80%
GPQA
—
Context 200KPrice $0.30 / $1.2 /M in·outSpeed 80 t/sCoding · Best value · Agents
KI
KimiOpen

Kimi K2.6

Huge document analysis and Chinese/English long-context apps.

Details
Arena
—
AA Index
43
SWE-bench
—
GPQA
—
Context 2MPrice $0.50 / $2 /M in·outSpeed 70 t/sLong context · Best value · Writing
GO
Google

Gemini 3 Flash

High-throughput apps, classification, and cost-sensitive Google stack workloads.

Details
Arena
1473
AA Index
42
SWE-bench
78%
GPQA
90%
Context 1MPrice $0.30 / $2.5 /M in·outSpeed 180 t/sSpeed · Best value · Multimodal
MI
Mistral

Mistral Large 3

EU data residency needs and multilingual production apps.

Details
Arena
1410
AA Index
40
SWE-bench
72%
GPQA
82%
Context 256KPrice $2 / $6 /M in·outSpeed 100 t/sWriting · Best value · Multimodal
ME
MetaOpen

Llama 4 Maverick

Enterprise self-host, fine-tunes, and cost-controlled chat at scale.

Details
Arena
1380
AA Index
38
SWE-bench
—
GPQA
—
Context 1MPrice $0.19 / $0.49 /M in·outSpeed 120 t/sBest value · Multimodal · Speed

Model Gauge · built by Cursor Grok 4.5 High · curated snapshot, not a live leaderboard

Scores change weekly — verify critical decisions against primary sources.