Kimi K3Recent
MoonshotJul 22, 2026
- SWE-V
- 93.4/100
- GPQA
- —
- OSWorld
- —
Frontier AI · August 2026 snapshot
Live rankings across reasoning, coding and agentic benchmarks — with real API pricing and context windows. Filter to your workload, sort by the metric you care about, and put up to four models head-to-head.
Showing 38 of 38 models · sorted by Overall composite (high → low, no data sinks last)
MoonshotJul 22, 2026
DeepSeekAug 13, 2026
MoonshotApr 8, 2026
AnthropicJun 9, 2026
DeepSeekApr 24, 2026
OpenAIJul 9, 2026
GoogleDec 17, 2025
GoogleFeb 19, 2026
GoogleNov 18, 2025
xAIJul 16, 2026
AnthropicOct 15, 2025
MoonshotNov 6, 2025
Z.aiJun 13, 2026
xAIAug 14, 2026
Z.aiAug 14, 2026
MetaAug 10, 2026
QwenAug 5, 2026
AnthropicJul 24, 2026
OpenAIJul 9, 2026
OpenAIJul 9, 2026
MetaJul 9, 2026
AnthropicJun 30, 2026
MicrosoftJun 2, 2026
AnthropicMay 28, 2026
OpenAIApr 24, 2026
OpenAIApr 24, 2026
Z.aiFeb 11, 2026
AmazonDec 2, 2025
AmazonDec 2, 2025
AmazonDec 2, 2025
xAIDec 1, 2025
MiniMaxOct 27, 2025
GoogleAug 1, 2025
MistralAug 1, 2025
QwenJul 24, 2025
QwenJul 23, 2025
MistralJul 10, 2025
MetaApr 5, 2025
Official naming and lineups as published by each lab — from GPT‑5.6's Sol/Terra/Luna tiers to Claude's new Fable line and Meta's Muse family.
No single number tells you which model fits your workload. Score a coding agent on BrowseComp or a research assistant on Terminal-Bench and you learn nothing. These are the families we track and why each exists:
Expert science QA, hardened knowledge exams, and abstract puzzle-solving.
GPQA Diamondleader: GPT-5.6 Sol 93.6%
Graduate-level, Google-proof science questions written by PhD experts. A proxy for deep scientific reasoning beyond memorization.
MMLU-Proleader: DeepSeek-V4-Pro (0813) 87.5%
Hardened successor to MMLU across 14 domains with ten answer options; measures broad knowledge plus robust reasoning.
Humanity's Last Examleader: Kimi K2 Thinking 44.9%
Extremely difficult multi-disciplinary exam designed at the edge of human knowledge; frontier scores remain far from saturation.
ARC-AGI-2leader: Gemini 3.1 Pro 77.1%
Abstract-reasoning puzzles requiring novel rule induction from few examples. Designed to resist memorization and pattern-matching shortcuts.
Resolving real GitHub issues, long-horizon software engineering, fresh contest problems.
SWE-bench Verifiedleader: Kimi K3 93.4%
Human-validated subset of real GitHub issues that must be resolved end-to-end in live repositories. The industry-standard agentic coding metric.
SWE-bench Proleader: Claude Fable 5 80%
Harder, contamination-resistant split of SWE-bench with longer-horizon enterprise-style tasks. Top frontier models score well below Verified.
LiveCodeBenchleader: DeepSeek-V4-Pro (0813) 93.5%
Continuously refreshed competitive-programming problems published after model cutoffs, neutralizing training-set contamination.
Terminal work, full-desktop computer control, and persistent web research.
Terminal-Bench 2leader: Grok 4.5 83.3%
Autonomous agent tasks executed inside a real terminal: build systems, debugging, sysadmin work across many turns of tool use.
OSWorld 2.0leader: GPT-5.6 Sol 62.6%
Computer-use agent benchmark driving full desktop environments — clicking, typing, cross-app workflows like a human operator.
BrowseCompleader: GPT-5.6 Sol 92.2%
Agentic web research: locating hard-to-find facts across the live internet through persistent search-and-read behavior.
Exact competition mathematics under strict answer checking.
AIME 2026leader: Kimi K2 Thinking 94.5%
American Invitational Mathematics Examination problems — competition math demanding exact multi-step symbolic reasoning.
Read the fine print