Frontier BenchAI model benchmarks · Jul 2026
ModelsCompareLabsMethodology
All models
GD

Gemini 3.1 Pro

Google DeepMind

Undisputed long-context leader at 2M tokens. Best $/token for batch reasoning.

FrontierProprietary2M contextReleased Mar 2026textvisionaudio-inaudio-outvideocode

API price · per 1M tokens

$2/$12

input / output

≈ $4.50/M blended

Benchmarks

Bar shows this model's score; tick shows the best score across all tracked models for context.

SWE-bench VerifiedReal GitHub issues, human-verified patches. The standard coding-agent benchmark.
80.6/ best 90.1
SWE-bench Pro2,294 harder real-world GitHub issues. Where the frontier separates now.
54.2/ best 69.2
Terminal-Bench 2.xAgentic command-line workflows — planning, iteration, tool coordination.
70.3/ best 84.0
GPQA DiamondGoogle-proof graduate-level science Q&A. General reasoning proxy.
94.3/ best 94.4
MMLU-ProMulti-task language understanding across 120+ academic subjects.
87.0/ best 93.0
HumanEval+Functional code generation from docstrings.
91.7/ best 96.0
MATH-500Competition-style math problems.
90.0/ best 97.8
Humanity's Last ExamHardest reasoning benchmark (with tools). ~50% = near frontier.
51.4/ best 57.9
AIME 2025American Invitational Mathematics Examination.
88.9/ best 96.5
OSWorldReal computer-use tasks across desktop apps.
76.2/ best 83.4

Other Google DeepMind models

GD

Gemini 3.5 Flash

Fast / Cheap

≈$3.38/M

Built by GLM-5.2 · Scores are directional, drawn from public mid-2026 benchmarks.

SWE-bench, GPQA, MMLU, AIME, HLE etc. are trademarks of their respective owners.