MODEL
ATLAS
ModelsBenchmarksBriefing

THE FRONTIER, MADE LEGIBLE

Find the model
for the work.

Compare capability, cost, and practical fit across the models shaping what’s next.

Live index · Updated 15 Jul 2026
7 MODELSScores normalized across available evaluations
✦

Google DeepMind

Gemini 3.1 Pro

Frontier1M ctx
Reasoning
96
Code
83
Vision
91
Speed
67
$2 / $12
⌘

OpenAI

GPT-5.5

Frontier400K ctx
Reasoning
95
Code
86
Vision
89
Speed
75
$1.75 / $14
A

Anthropic

Claude Opus 4.7

Frontier1M ctx
Reasoning
94
Code
89
Vision
84
Speed
58
$5 / $25
𝕏

xAI

Grok 4

Frontier256K ctx
Reasoning
91
Code
79
Vision
82
Speed
72
$3 / $15
Q

Alibaba

Qwen3-Max

API262K ctx
Reasoning
85
Code
80
Vision
79
Speed
80
$1.20 / $6
M

Mistral AI

Mistral Large 3

Open weights256K ctx
Reasoning
80
Code
76
Vision
77
Speed
91
$0.50 / $1.50
∞

Meta

Llama 4 Maverick

Open weights10M ctx
Reasoning
79
Code
72
Vision
80
Speed
88
Self-host
AT A GLANCE
LEADING OVERALL

Gemini 3.1 Pro

Highest composite score

✦
BEST FOR CODE

Claude Opus 4.7

Top agentic coding score

A
BEST VALUE

Mistral Large 3

High capability, low cost

M

Benchmark reality check Scores are not directly comparable across every lab. Open a model to see its evaluation conditions.

How we read benchmarks

THE WEEKLY READOUT

Model movement,
without the noise.

New releases, benchmark shifts, and what they mean for teams choosing models.

2 selectedCompare models