Frontier AI · August 2026 snapshot

Every frontier model, measured on what matters

Live rankings across reasoning, coding and agentic benchmarks — with real API pricing and context windows. Filter to your workload, sort by the metric you care about, and put up to four models head-to-head.

38
models tracked
13
labs
11
benchmarks
16
open-weight
Open the boardCompare models

Showing 38 of 38 models · sorted by Overall composite (high → low, no data sinks last)

N/AK

Kimi K3Recent

Moonshot·Jul 22, 2026

Open
SWE-V
93.4/100
GPQA
—
OSWorld
—
  • $3/$15
  • ·
  • 1M ctx
  • ·
  • Open weights
#1D

DeepSeek-V4-Pro (0813)New

DeepSeek·Aug 13, 2026

Open
SWE-V
80.6/100
GPQA
90.1/100
OSWorld
—
  • $0.7/$3.48
  • ·
  • 1M ctx
  • ·
  • Open · MIT
N/AK

Kimi K2.6

Moonshot·Apr 8, 2026

Open
SWE-V
80.2/100
GPQA
—
OSWorld
—
  • $0.6/$2.50
  • ·
  • 256K ctx
  • ·
  • Open weights
N/AA

Claude Fable 5Recent

Anthropic·Jun 9, 2026

SWE-Pro
80/100
  • $8/$40
  • ·
  • 1M ctx
  • ·
  • API
N/AD

DeepSeek-V4 Flash

DeepSeek·Apr 24, 2026

Open
SWE-V
79/100
GPQA
—
OSWorld
—
  • $0.42/$1.98
  • ·
  • 512K ctx
  • ·
  • Open · MIT
#2O

GPT-5.6 SolRecent

OpenAI·Jul 9, 2026

SWE-V
—
GPQA
93.6/100
OSWorld
62.6/100
  • $5/$30
  • ·
  • 1.1M ctx
  • ·
  • API
N/AG

Gemini 3 Flash

Google·Dec 17, 2025

SWE-V
78/100
GPQA
—
OSWorld
—
  • $0.5/$3
  • ·
  • 1M ctx
  • ·
  • API
N/AG

Gemini 3.1 Pro

Google·Feb 19, 2026

ARC-AGI-2
77.1/100
  • $2/$12
  • ·
  • 1M ctx
  • ·
  • API
N/AG

Gemini 3 Pro

Google·Nov 18, 2025

SWE-V
76.2/100
GPQA
—
OSWorld
—
  • $2/$12
  • ·
  • 1M ctx
  • ·
  • API
N/A𝕏

Grok 4.5Recent

xAI·Jul 16, 2026

SWE-Pro
64.7/100
Term-Bench
83.3/100
  • $2/$6
  • ·
  • 500K ctx
  • ·
  • API
N/AA

Claude Haiku 4.5

Anthropic·Oct 15, 2025

SWE-V
73.3/100
GPQA
—
OSWorld
—
  • $1/$5
  • ·
  • 200K ctx
  • ·
  • API
#3K

Kimi K2 Thinking

Moonshot·Nov 6, 2025

Open
SWE-V
71.3/100
GPQA
—
OSWorld
—
  • $0.6/$2.50
  • ·
  • 256K ctx
  • ·
  • Open weights
N/AZ

GLM-5.2Recent

Z.ai·Jun 13, 2026

Open
SWE-Pro
62.1/100
  • $1.40/$4.40
  • ·
  • 1M ctx
  • ·
  • Open · MIT
N/A𝕏

Grok 4.6New

xAI·Aug 14, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • $2/$6
  • ·
  • 500K ctx
  • ·
  • API
N/AZ

GLM-5.3New

Z.ai·Aug 14, 2026

OpenSub first
SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 1M ctx
  • ·
  • Open weights
N/AM

Muse Glimmer 30BNew

Meta·Aug 10, 2026

Open
SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 256K ctx
  • ·
  • Open weights
N/AQ

Qwen3.8 MaxNew

Qwen·Aug 5, 2026

Open
SWE-V
—
GPQA
—
OSWorld
—
  • $2/$6
  • ·
  • 1M ctx
  • ·
  • Open weights
N/AA

Claude Opus 5Recent

Anthropic·Jul 24, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • $5/$25
  • ·
  • 200K ctx
  • ·
  • API
N/AO

GPT-5.6 TerraRecent

OpenAI·Jul 9, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • $2/$12
  • ·
  • 1.1M ctx
  • ·
  • API
N/AO

GPT-5.6 LunaRecent

OpenAI·Jul 9, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • $0.2/$1.20
  • ·
  • 1.1M ctx
  • ·
  • API
N/AM

Muse Spark 1.1Recent

Meta·Jul 9, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 500K ctx
  • ·
  • API
N/AA

Claude Sonnet 5Recent

Anthropic·Jun 30, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • $3/$15
  • ·
  • 1M ctx
  • ·
  • API
N/AMS

MAI-Thinking-1Recent

Microsoft·Jun 2, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 256K ctx
  • ·
  • API
N/AA

Claude Opus 4.8Recent

Anthropic·May 28, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • $5/$25
  • ·
  • 200K ctx
  • ·
  • API
N/AO

GPT-5.5 Pro

OpenAI·Apr 24, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 400K ctx
  • ·
  • API
N/AO

GPT-5.5

OpenAI·Apr 24, 2026

SWE-V
—
GPQA
—
OSWorld
—
  • $1.25/$10
  • ·
  • 400K ctx
  • ·
  • API
N/AZ

GLM-5

Z.ai·Feb 11, 2026

Open
SWE-V
—
GPQA
—
OSWorld
—
  • $1/$3.20
  • ·
  • 200K ctx
  • ·
  • Open · MIT
N/Aa

Nova 2 Pro

Amazon·Dec 2, 2025

Preview
SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 300K ctx
  • ·
  • API
N/Aa

Nova 2 Lite

Amazon·Dec 2, 2025

SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 250K ctx
  • ·
  • API
N/Aa

Nova 2 Omni

Amazon·Dec 2, 2025

Preview
SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 300K ctx
  • ·
  • API
N/A𝕏

Grok 4 Fast Reasoning

xAI·Dec 1, 2025

Open
SWE-V
—
GPQA
—
OSWorld
—
  • $0.2/$0.8
  • ·
  • 2M ctx
  • ·
  • Open weights
N/AmM

MiniMax-M2

MiniMax·Oct 27, 2025

Open
SWE-V
—
GPQA
—
OSWorld
—
  • $0.3/$1.20
  • ·
  • 205K ctx
  • ·
  • Open · MIT
N/AG

Gemini 3 Deep Think

Google·Aug 1, 2025

Sub first
SWE-V
—
GPQA
—
OSWorld
—
  • —/—
  • ·
  • 1M ctx
  • ·
  • API
N/AMi

Mistral Medium 3.1

Mistral·Aug 1, 2025

SWE-V
—
GPQA
—
OSWorld
—
  • $0.4/$2
  • ·
  • 128K ctx
  • ·
  • API
N/AQ

Qwen3 235B Instruct 2507

Qwen·Jul 24, 2025

Open
SWE-V
—
GPQA
—
OSWorld
—
  • $0.22/$0.88
  • ·
  • 262K ctx
  • ·
  • Open · Apache 2.0
N/AQ

Qwen3 Coder Flash

Qwen·Jul 23, 2025

Open
SWE-V
—
GPQA
—
OSWorld
—
  • $0.3/$1.20
  • ·
  • 262K ctx
  • ·
  • Open · Apache 2.0
N/AMi

Devstral Medium 2507

Mistral·Jul 10, 2025

Open
SWE-V
—
GPQA
—
OSWorld
—
  • $0.4/$2
  • ·
  • 128K ctx
  • ·
  • Open weights
N/AM

Llama 4 Maverick

Meta·Apr 5, 2025

Open
SWE-V
—
GPQA
—
OSWorld
—
  • $0.27/$0.85
  • ·
  • 1M ctx
  • ·
  • Open weights
ModelOverallSWE-VGPQAOSWorld$ & ctxReleasedAdd to compare
N/AK
Kimi K3Moonshot
93.493.4——$6/1MJul 22, 2026
#1D
DeepSeek-V4-Pro (0813)DeepSeek · new
87.980.690.1—$1.40/1MAug 13, 2026
N/AK
Kimi K2.6Moonshot
80.280.2——$1.07/256KApr 8, 2026
N/AA
Claude Fable 5Anthropic
80———$16/1MJun 9, 2026
N/AD
DeepSeek-V4 FlashDeepSeek
7979——$0.81/512KApr 24, 2026
#2O
GPT-5.6 SolOpenAI
78.3—93.662.6$11.25/1.1MJul 9, 2026
N/AG
Gemini 3 FlashGoogle
7878——$1.13/1MDec 17, 2025
N/AG
Gemini 3.1 ProGoogle
77.1———$4.50/1MFeb 19, 2026
N/AG
Gemini 3 ProGoogle
76.276.2——$4.50/1MNov 18, 2025
N/A𝕏
Grok 4.5xAI
74———$3/500KJul 16, 2026
N/AA
Claude Haiku 4.5Anthropic
73.373.3——$2/200KOct 15, 2025
#3K
Kimi K2 ThinkingMoonshot
70.271.3——$1.07/256KNov 6, 2025
N/AZ
GLM-5.2Z.ai
62.1———$2.15/1MJun 13, 2026
N/A𝕏
Grok 4.6xAI · new
————$3/500KAug 14, 2026
N/AZ
GLM-5.3Z.ai · new
—————/1MAug 14, 2026
N/AM
Muse Glimmer 30BMeta · new
—————/256KAug 10, 2026
N/AQ
Qwen3.8 MaxQwen · new
————$3/1MAug 5, 2026
N/AA
Claude Opus 5Anthropic
————$10/200KJul 24, 2026
N/AO
GPT-5.6 TerraOpenAI
————$4.50/1.1MJul 9, 2026
N/AO
GPT-5.6 LunaOpenAI
————$0.45/1.1MJul 9, 2026
N/AM
Muse Spark 1.1Meta
—————/500KJul 9, 2026
N/AA
Claude Sonnet 5Anthropic
————$6/1MJun 30, 2026
N/AMS
MAI-Thinking-1Microsoft
—————/256KJun 2, 2026
N/AA
Claude Opus 4.8Anthropic
————$10/200KMay 28, 2026
N/AO
GPT-5.5 ProOpenAI
—————/400KApr 24, 2026
N/AO
GPT-5.5OpenAI
————$3.44/400KApr 24, 2026
N/AZ
GLM-5Z.ai
————$1.55/200KFeb 11, 2026
N/Aa
Nova 2 ProAmazon
—————/300KDec 2, 2025
N/Aa
Nova 2 LiteAmazon
—————/250KDec 2, 2025
N/Aa
Nova 2 OmniAmazon
—————/300KDec 2, 2025
N/A𝕏
Grok 4 Fast ReasoningxAI
————$0.35/2MDec 1, 2025
N/AmM
MiniMax-M2MiniMax
————$0.52/205KOct 27, 2025
N/AG
Gemini 3 Deep ThinkGoogle
—————/1MAug 1, 2025
N/AMi
Mistral Medium 3.1Mistral
————$0.80/128KAug 1, 2025
N/AQ
Qwen3 235B Instruct 2507Qwen
————$0.39/262KJul 24, 2025
N/AQ
Qwen3 Coder FlashQwen
————$0.52/262KJul 23, 2025
N/AMi
Devstral Medium 2507Mistral
————$0.80/128KJul 10, 2025
N/AM
Llama 4 MaverickMeta
————$0.42/1MApr 5, 2025

Just landed ≤ 30 days

  • 𝕏Grok 4.6New · 12d ago
  • ZGLM-5.3New · 12d ago
  • DDeepSeek-V4-Pro (0813)New · 13d ago
  • MMuse Glimmer 30BNew · 16d ago
  • QQwen3.8 MaxNew · 21d ago

Labs covered

Official naming and lineups as published by each lab — from GPT‑5.6's Sol/Terra/Luna tiers to Claude's new Fable line and Meta's Muse family.

  • OOpenAI
    San Francisco, US5 models
  • AAnthropic
    San Francisco, US5 models
  • GGoogle DeepMind
    Mountain View, US4 models
  • 𝕏xAI
    Palo Alto, US3 models1 open
  • KMoonshot AI
    Beijing, CN3 models3 open
  • ZZhipu AI (Z.ai)
    Beijing, CN3 models3 open
  • DDeepSeek
    Hangzhou, CN2 models2 open
  • QAlibaba (Qwen)
    Hangzhou, CN3 models3 open
  • MMeta Superintelligence Labs
    Menlo Park, US3 models2 open
  • MSMicrosoft AI
    Redmond, US1 model
  • aAmazon AGI
    Seattle, US3 models
  • MiMistral AI
    Paris, FR2 models1 open
  • mMMiniMax
    Shanghai, CN1 model1 open

What the benchmarks actually measure

No single number tells you which model fits your workload. Score a coding agent on BrowseComp or a research assistant on Terminal-Bench and you learn nothing. These are the families we track and why each exists:

Reasoning & knowledge

Expert science QA, hardened knowledge exams, and abstract puzzle-solving.

  • GPQA Diamondleader: GPT-5.6 Sol 93.6%

    Graduate-level, Google-proof science questions written by PhD experts. A proxy for deep scientific reasoning beyond memorization.

  • MMLU-Proleader: DeepSeek-V4-Pro (0813) 87.5%

    Hardened successor to MMLU across 14 domains with ten answer options; measures broad knowledge plus robust reasoning.

  • Humanity's Last Examleader: Kimi K2 Thinking 44.9%

    Extremely difficult multi-disciplinary exam designed at the edge of human knowledge; frontier scores remain far from saturation.

  • ARC-AGI-2leader: Gemini 3.1 Pro 77.1%

    Abstract-reasoning puzzles requiring novel rule induction from few examples. Designed to resist memorization and pattern-matching shortcuts.

Coding

Resolving real GitHub issues, long-horizon software engineering, fresh contest problems.

  • SWE-bench Verifiedleader: Kimi K3 93.4%

    Human-validated subset of real GitHub issues that must be resolved end-to-end in live repositories. The industry-standard agentic coding metric.

  • SWE-bench Proleader: Claude Fable 5 80%

    Harder, contamination-resistant split of SWE-bench with longer-horizon enterprise-style tasks. Top frontier models score well below Verified.

  • LiveCodeBenchleader: DeepSeek-V4-Pro (0813) 93.5%

    Continuously refreshed competitive-programming problems published after model cutoffs, neutralizing training-set contamination.

Agents & computer use

Terminal work, full-desktop computer control, and persistent web research.

  • Terminal-Bench 2leader: Grok 4.5 83.3%

    Autonomous agent tasks executed inside a real terminal: build systems, debugging, sysadmin work across many turns of tool use.

  • OSWorld 2.0leader: GPT-5.6 Sol 62.6%

    Computer-use agent benchmark driving full desktop environments — clicking, typing, cross-app workflows like a human operator.

  • BrowseCompleader: GPT-5.6 Sol 92.2%

    Agentic web research: locating hard-to-find facts across the live internet through persistent search-and-read behavior.

Math

Exact competition mathematics under strict answer checking.

  • AIME 2026leader: Kimi K2 Thinking 94.5%

    American Invitational Mathematics Examination problems — competition math demanding exact multi-step symbolic reasoning.

Read the fine print

  • Missing scores show as “—”. We never extrapolate numbers a lab hasn't published; sparse cells mean exactly that.
  • Many headline figures are vendor-reported on vendor-chosen harnesses. Where independent harnesses disagree materially (e.g. Kimi K3, DeepSeek V4 SWE-bench runs) the caveat is noted on the model page.
  • Benchmark columns aren't perfectly comparable across labs — SWE-bench Verified vs Pro vs Marathon measure different difficulty ceilings, so within-column comparison only.
  • Prices are standard-tier USD per million tokens from public rate cards; subscription-first releases (GLM-5.3, Deep Think) have none.
  • Release dates reflect best-documented announcements; some sources differ by days. Data compiled August 26, 2026.