GPT-5.6 Sol
OpenAI's working flagship — elite math agency plus broad tooling.
ModelLens
AI benchmarks, decoded
Frontier models · Sep 2026 edition
We track flagship models from every major lab across knowledge, reasoning, coding and multimodal benchmarks — then help you filter by your job to be done, not just the hype score.
I want a model for…
Showing all models, ranked by blended overall score.
13 of 13 models · sorted by Overall score
GPT-5.6 Sol
OpenAI · 78.3 Overall score
Grok 4.6
SpaceXAI · 77.6 Overall score
Claude Fable 5.1
Anthropic · 72.7 Overall score
OpenAI's working flagship — elite math agency plus broad tooling.
Ties Sol on the composite at one-fifth the output price.
Released yesterday — tops the AA Index and leads SWE-Pro and HLE.
The analyst's Flash — best published business-document score.
Near-Fable capability at half the running cost, per Anthropic.
1.6T open MoE built for long-horizon agent workloads.
Dense 128B open flagship — chat, reasoning and code in one.
Largest open weights ever — 2.8T MoE with native video.
Released today — leads the 24-model DeepSWE board.
Released today — Google's best reasoning/coding at Flash price.
Superseded within a month — check Spark 1.3 first.
Released today — 2.4T open MoE, #1 on Code Arena WebDev.
Announced Sep 1 — first 'Critical' cyber-capability rating; gated rollout.
Overall (0–100) = mean across all 8 benchmarks of each score relative to the best in this lineup, with a neutral 50 for unreported cells — so breadth of evidence counts and thin day-zero rows can't top the board on one number. The bench count next to every Overall (e.g. 6/8) tells you how much evidence backs it. All figures are vendor/lab-reported launch numbers compiled Sep 2, 2026; independent re-runs sometimes differ by several points (Epoch AI put Fable 5 at GPQA 85.9 vs 92.6 claimed, V4-Pro SWE-Verified at 77.6 vs 80.6). Prices are public list API rates per 1M tokens; “—” means undisclosed, not zero.
AA Intelligence Index Composite
Artificial Analysis 9-bench composite, 0–100. The closest thing to one single leaderboard.
SWE-bench Verified Coding
500 real GitHub issues. Grok's 95.6 is lab-reported; Epoch AI re-ran V4-Pro at 77.6 vs 80.6 claimed.
SWE-bench Pro Coding
731 harder production tasks from copyleft repos. The coding bench labs still put in launch posts.
GPQA Diamond Reasoning
198 PhD science questions. Effectively saturated — reporters bunch within ~5pts; Anthropic will stop reporting it.
Humanity's Last Exam Knowledge
3k hardest academic questions. Methods differ by vendor (tools vs no-tools) — compare loosely.
DeepSWE v1.1 Coding
Long-horizon end-to-end software engineering. Google claims 3.8 Flash leads but published no %.
ARC-AGI-2 Reasoning
Visual abstraction puzzles via the ARC Prize verified board. Gaps mean no submission, not a zero.
AA-AnalystAgent Agentic
Real spreadsheets + business documents (pass^5). Fable 5.1 ran a fallback config here.
Terminal-Bench is intentionally not a column: vendors report incompatible versions (v2.1 vs v3.0 vs v4.0 vs Science split). Notable rows live in the model notes instead. Brand colors are indicative dots, not official logos. Model availability and pricing change fast — verify with the lab before building.