Frontier BenchAI model benchmarks · Jul 2026
ModelsCompareLabsMethodology
Updated Jul 2026 · 20 models · 9 labs

Which AI model should you actually use?

Compare frontier and open-weight models side by side. Filter by lab, license, modality, and price. Sort by the benchmark that matters to you. Pin a few and head to the comparison view.

Open comparison viewHow benchmarks work

At a glance

Top 3 per benchmark · tap a card for detail

SWE-bench Verified

Real GitHub issues, human-verified patches. The standard coding-agent benchmark.

  1. 1OGPT-5.5 Pro90.1
  2. 2OGPT-5.588.7
  3. 3AClaude Opus 4.888.6

SWE-bench Pro

2,294 harder real-world GitHub issues. Where the frontier separates now.

  1. 1AClaude Opus 4.869.2
  2. 2AClaude Opus 4.764.3
  3. 3OGPT-5.5 Pro63.0

GPQA Diamond

Google-proof graduate-level science Q&A. General reasoning proxy.

  1. 1OGPT-5.5 Pro94.4
  2. 2GDGemini 3.1 Pro94.3
  3. 3AClaude Opus 4.794.2

MMLU-Pro

Multi-task language understanding across 120+ academic subjects.

  1. 1OGPT-5.5 Pro93.0
  2. 2OGPT-5.592.4
  3. 3GDGemini 3.1 Pro87.0

Terminal-Bench 2.x

Agentic command-line workflows — planning, iteration, tool coordination.

  1. 1OGPT-5.5 Pro84.0
  2. 2OGPT-5.582.7
  3. 3GDGemini 3.5 Flash76.2

MATH-500

Competition-style math problems.

  1. 1OGPT-5.5 Pro97.8
  2. 2OGPT-5.597.1
  3. 3XGrok 4.2096.5

Best value

Highest average score per dollar (blended price)

  1. 1MAMistral Small 4≈$0.15/M
  2. 2DDeepSeek V4 Flash≈$0.18/M
  3. 3MALlama 4 Scout≈$0.22/M
TierModalities
Showing 20 of 20 models
A

Claude Opus 4.8

Anthropic

Current agentic-coding and computer-use leader. Best default for engineering teams.

FrontierProprietary1M ctxlong ctxagents
SWE-bench Verified88.6
SWE-bench Pro69.2
GPQA Diamond93.6
MMLU-Pro84.8

Per 1M tok

$5 / $25

≈$10.00/MDetails
GD

Gemini 3.5 Flash

Google DeepMind

Beats the prior Pro flagship on coding at 4× the speed. Volume-pipeline darling.

Fast / CheapProprietary1M ctxlong ctxmultimodalagents
SWE-bench Verified79.0
SWE-bench Pro58.0
GPQA Diamond91.0
MMLU-Pro85.0

Per 1M tok

$1.5 / $9

≈$3.38/MDetails
D

DeepSeek V4 Pro

DeepSeek

Best open-weight coder. MIT license, 1M context, 49B active MoE.

Open-weightMIT1M ctxlong ctx
SWE-bench Verified83.7
SWE-bench Pro55.4
GPQA Diamond85.0
MMLU-Pro82.0

Per 1M tok

$0.435 / $0.87

≈$0.54/MDetails
D

DeepSeek V4 Flash

DeepSeek

Cheapest credible API on the market. Matches reasoning with a larger thinking budget.

Open-weightMIT1M ctx
SWE-bench Verified72.0
SWE-bench Pro—
GPQA Diamond76.0
MMLU-Pro75.0

Per 1M tok

$0.14 / $0.28

≈$0.18/MDetails
MA

Mistral Small 4

Mistral AI

119B dense. Runs on a single 8×H100 node — easy self-host generalist.

Open-weightApache 2.0128K ctx
SWE-bench Verified58.0
SWE-bench Pro—
GPQA Diamond70.0
MMLU-Pro72.0

Per 1M tok

$0.1 / $0.3

≈$0.15/MDetails
O

GPT-5.5

OpenAI

Best terminal/CLI autonomy and dense academic reasoning. Effort variants up to max.

FrontierProprietary400K ctxmultimodalagents
SWE-bench Verified88.7
SWE-bench Pro58.6
GPQA Diamond93.6
MMLU-Pro92.4

Per 1M tok

$10 / $30

≈$15.00/MDetails
O

GPT-5.5 Pro

OpenAI

Higher-effort Pro tier. Wins on HLE with tools and FrontierMath.

FrontierProprietary400K ctxagents
SWE-bench Verified90.1
SWE-bench Pro63.0
GPQA Diamond94.4
MMLU-Pro93.0

Per 1M tok

$15 / $60

≈$26.25/MDetails
AQ

Qwen 3.7 Max

Alibaba (Qwen)

Apache-2.0 hosted frontier model. Top multilingual and agentic-coding open option.

FrontierApache 2.01M ctxagents
SWE-bench Verified81.0
SWE-bench Pro60.6
GPQA Diamond88.4
MMLU-Pro84.0

Per 1M tok

$2.5 / $7.5

≈$3.75/MDetails
GD

Gemini 3.1 Pro

Google DeepMind

Undisputed long-context leader at 2M tokens. Best $/token for batch reasoning.

FrontierProprietary2M ctxlong ctxmultimodal
SWE-bench Verified80.6
SWE-bench Pro54.2
GPQA Diamond94.3
MMLU-Pro87.0

Per 1M tok

$2 / $12

≈$4.50/MDetails
A

Claude Opus 4.7

Anthropic

Previous flagship. Still a strong repository-reasoning pick; superseded by 4.8.

FrontierProprietary1M ctxlong ctx
SWE-bench Verified87.6
SWE-bench Pro64.3
GPQA Diamond94.2
MMLU-Pro84.0

Per 1M tok

$5 / $25

≈$10.00/MDetails
X

Grok 4.20

xAI

2M context at frontier-cheap output pricing. Strong on raw math and HumanEval.

FrontierProprietary2M ctxlong ctx
SWE-bench Verified78.0
SWE-bench Pro56.0
GPQA Diamond74.5
MMLU-Pro80.0

Per 1M tok

$2 / $6

≈$3.00/MDetails
AQ

Qwen 3.5 Coder

Alibaba (Qwen)

Code-tuned open model with a huge fine-tune ecosystem. IDE autocomplete favorite.

Open-weightApache 2.0256K ctx
SWE-bench Verified75.0
SWE-bench Pro48.0
GPQA Diamond—
MMLU-Pro—

Per 1M tok

$0.2 / $0.6

≈$0.30/MDetails
A

Claude Sonnet 4.7

Anthropic

Best price-to-capability ratio in the Claude lineup for everyday coding.

MidProprietary1M ctxlong ctx
SWE-bench Verified77.5
SWE-bench Pro51.0
GPQA Diamond83.5
MMLU-Pro78.0

Per 1M tok

$3 / $15

≈$6.00/MDetails
A

Claude Haiku 4.7

Anthropic

Sub-second latency tier for chat and classification at scale.

Fast / CheapProprietary250K ctx
SWE-bench Verified60.0
SWE-bench Pro—
GPQA Diamond65.0
MMLU-Pro70.0

Per 1M tok

$0.8 / $4

≈$1.60/MDetails
X

Grok 4.1 Fast

xAI

Budget-tier coding leader. Cheapest credible agent for high-volume workloads.

Fast / CheapProprietary2M ctx
SWE-bench Verified70.0
SWE-bench Pro—
GPQA Diamond68.0
MMLU-Pro76.0

Per 1M tok

$0.2 / $0.5

≈$0.28/MDetails
MA

Llama 4 Maverick

Meta AI

MoE flagship with deep ecosystem tooling. 17B active, 400B total.

Open-weightLlama license1M ctxlong ctxmultimodal
SWE-bench Verified70.0
SWE-bench Pro44.0
GPQA Diamond80.0
MMLU-Pro80.5

Per 1M tok

$0.2 / $0.6

≈$0.30/MDetails
MA

Llama 4 Scout

Meta AI

10M-token context — the only realistic choice for whole-repo or whole-corpus input.

Open-weightLlama license10M ctxlong ctxmultimodal
SWE-bench Verified55.0
SWE-bench Pro—
GPQA Diamond72.0
MMLU-Pro75.0

Per 1M tok

$0.15 / $0.45

≈$0.22/MDetails
MA

Mistral Large 3

Mistral AI

675B dense flagship. Cleanest license for EU data-sovereign deployment.

Open-weightApache 2.0256K ctx
SWE-bench Verified68.0
SWE-bench Pro—
GPQA Diamond78.0
MMLU-Pro78.0

Per 1M tok

$0.4 / $1.2

≈$0.60/MDetails
O

GPT-5.1 Mini

OpenAI

Workhorse small model — cheap, fast, multimodal, ubiquitous.

Fast / CheapProprietary200K ctxmultimodal
SWE-bench Verified70.0
SWE-bench Pro—
GPQA Diamond75.0
MMLU-Pro80.0

Per 1M tok

$0.4 / $1.6

≈$0.70/MDetails
ZA

GLM-4.6

Zhipu AI (Z.ai)

Cost-effective coding & agentic model. Strong math for the price.

MidOpen weights200K ctxagents
SWE-bench Verified68.0
SWE-bench Pro—
GPQA Diamond81.0
MMLU-Pro82.9

Per 1M tok

$0.43 / $1.74

≈$0.76/MDetails

Built by GLM-5.2 · Scores are directional, drawn from public mid-2026 benchmarks.

SWE-bench, GPQA, MMLU, AIME, HLE etc. are trademarks of their respective owners.