◍

ModelLens

AI benchmarks, decoded

● Updated Sep 202613 models · 9 labs · 8 benchmarks
⌕

Frontier models · Sep 2026 edition

Which AI model should you actually use?

We track flagship models from every major lab across knowledge, reasoning, coding and multimodal benchmarks — then help you filter by your job to be done, not just the hype score.

Models tracked
13
9 labs covered
Top overall
GPT-5.6 Sol
78.3 overall
Best coder
Grok 4.6
95.6 SWE-Verified
Cheapest open
Mistral Medium 3.5
$1.5 / 1M in

I want a model for…

Showing all models, ranked by blended overall score.

13 of 13 models · sorted by Overall score

  1. 🥇

    GPT-5.6 Sol

    OpenAI · 78.3 Overall score

  2. 🥈

    Grok 4.6

    SpaceXAI · 77.6 Overall score

  3. 🥉

    Claude Fable 5.1

    Anthropic · 72.7 Overall score

OpenAI#1

GPT-5.6 Sol

OpenAI's working flagship — elite math agency plus broad tooling.

CodingReasoningAgents
Overall 78.3 · 6/8 benches$5→$30 · —
AA Index (composite)61.0
SWE-bench Verified74.9
SWE-bench Pro64.6
SpaceXAI#2

Grok 4.6

Ties Sol on the composite at one-fifth the output price.

BudgetReasoningAgents
Overall 77.6 · 6/8 benches$2→$6 · 500K
AA Index (composite)61.0
SWE-bench Verified95.6
SWE-bench Pro—
Anthropic#3NEW

Claude Fable 5.1

Released yesterday — tops the AA Index and leads SWE-Pro and HLE.

CodingAgentsEnterprise
Overall 72.7 · 4/8 benches$10→$50 · 1M
AA Index (composite)66.0
SWE-bench Verified—
SWE-bench Pro81.2
Google DeepMind#4

Gemini 3.7 Flash

The analyst's Flash — best published business-document score.

BudgetMultimodalAgents
Overall 72.0 · 4/8 benchesprice TBC
AA Index (composite)56.0
SWE-bench Verified—
SWE-bench Pro—
Anthropic#5

Claude Opus 5

Near-Fable capability at half the running cost, per Anthropic.

CodingAgentsEnterprise
Overall 65.0 · 3/8 benches$5→$25 · 1M
AA Index (composite)63.0
SWE-bench Verified80.8
SWE-bench Pro—
DeepSeek#6OPEN

DeepSeek-V4-Pro

1.6T open MoE built for long-horizon agent workloads.

CodingOpenReasoning
Overall 64.2 · 4/8 benchesprice TBC
AA Index (composite)—
SWE-bench Verified80.6
SWE-bench Pro55.4
Mistral AI#7OPEN

Mistral Medium 3.5

Dense 128B open flagship — chat, reasoning and code in one.

OpenCodingEnterprise
Overall 57.5 · 2/8 benches$1.5→$7.5 · 256K
AA Index (composite)—
SWE-bench Verified77.6
SWE-bench Pro—
Moonshot AI#8OPEN

Kimi K3

Largest open weights ever — 2.8T MoE with native video.

OpenLong contextMultimodal
Overall 56.4 · 2/8 benches$3→$15 · 1M
AA Index (composite)57.0
SWE-bench Verified—
SWE-bench Pro—
Meta#9NEW

Muse Spark 1.3

Released today — leads the 24-model DeepSWE board.

CodingAgents
Overall 56.3 · 1/8 benchesprice TBC
AA Index (composite)—
SWE-bench Verified—
SWE-bench Pro—
Google DeepMind#10NEW

Gemini 3.8 Flash

Released today — Google's best reasoning/coding at Flash price.

BudgetCodingAgents
Overall 55.0 · 1/8 benches$0.75→$3.75 · —
AA Index (composite)—
SWE-bench Verified—
SWE-bench Pro—
Meta#11

Muse Spark 1.2

Superseded within a month — check Spark 1.3 first.

Agents
Overall 54.5 · 1/8 benchesprice TBC
AA Index (composite)57.0
SWE-bench Verified—
SWE-bench Pro—
Alibaba (Qwen)#12OPENNEW

Qwen3.8-Max

Released today — 2.4T open MoE, #1 on Code Arena WebDev.

CodingOpenBudget
Overall 54.2 · 1/8 benches$2→$6 · 1M
AA Index (composite)—
SWE-bench Verified—
SWE-bench Pro67.7
OpenAI#13ANNOUNCEDNEW

Astra

Announced Sep 1 — first 'Critical' cyber-capability rating; gated rollout.

ReasoningAgents
Overall 50.0 · 0/8 benchesprice TBC
AA Index (composite)—
SWE-bench Verified—
SWE-bench Pro—

How we score · what each benchmark means

Overall (0–100) = mean across all 8 benchmarks of each score relative to the best in this lineup, with a neutral 50 for unreported cells — so breadth of evidence counts and thin day-zero rows can't top the board on one number. The bench count next to every Overall (e.g. 6/8) tells you how much evidence backs it. All figures are vendor/lab-reported launch numbers compiled Sep 2, 2026; independent re-runs sometimes differ by several points (Epoch AI put Fable 5 at GPQA 85.9 vs 92.6 claimed, V4-Pro SWE-Verified at 77.6 vs 80.6). Prices are public list API rates per 1M tokens; “—” means undisclosed, not zero.

AA Intelligence Index Composite

Artificial Analysis 9-bench composite, 0–100. The closest thing to one single leaderboard.

SWE-bench Verified Coding

500 real GitHub issues. Grok's 95.6 is lab-reported; Epoch AI re-ran V4-Pro at 77.6 vs 80.6 claimed.

SWE-bench Pro Coding

731 harder production tasks from copyleft repos. The coding bench labs still put in launch posts.

GPQA Diamond Reasoning

198 PhD science questions. Effectively saturated — reporters bunch within ~5pts; Anthropic will stop reporting it.

Humanity's Last Exam Knowledge

3k hardest academic questions. Methods differ by vendor (tools vs no-tools) — compare loosely.

DeepSWE v1.1 Coding

Long-horizon end-to-end software engineering. Google claims 3.8 Flash leads but published no %.

ARC-AGI-2 Reasoning

Visual abstraction puzzles via the ARC Prize verified board. Gaps mean no submission, not a zero.

AA-AnalystAgent Agentic

Real spreadsheets + business documents (pass^5). Fable 5.1 ran a fallback config here.

Terminal-Bench is intentionally not a column: vendors report incompatible versions (v2.1 vs v3.0 vs v4.0 vs Science split). Notable rows live in the model notes instead. Brand colors are indicative dots, not official logos. Model availability and pricing change fast — verify with the lab before building.

ModelLens — an independent benchmark explorer.

Data snapshot Sep 2, 2026 · 13 models · Built with Next.js + Tailwind (mobile-first) · Not affiliated with OpenAI, Anthropic, Google, Meta, SpaceXAI, Mistral, DeepSeek, Alibaba or Moonshot.