Frontier BenchAI model benchmarks · Jul 2026
ModelsCompareLabsMethodology
All models
O

GPT-5.5 Pro

OpenAI

Higher-effort Pro tier. Wins on HLE with tools and FrontierMath.

FrontierProprietary400K contextReleased Apr 2026textvisionaudio-inaudio-outcode

API price · per 1M tokens

$15/$60

input / output

≈ $26.25/M blended

Benchmarks

Bar shows this model's score; tick shows the best score across all tracked models for context.

SWE-bench VerifiedReal GitHub issues, human-verified patches. The standard coding-agent benchmark.
90.1/ best 90.1

Best in class

SWE-bench Pro2,294 harder real-world GitHub issues. Where the frontier separates now.
63.0/ best 69.2
Terminal-Bench 2.xAgentic command-line workflows — planning, iteration, tool coordination.
84.0/ best 84.0

Best in class

GPQA DiamondGoogle-proof graduate-level science Q&A. General reasoning proxy.
94.4/ best 94.4

Best in class

MMLU-ProMulti-task language understanding across 120+ academic subjects.
93.0/ best 93.0

Best in class

HumanEval+Functional code generation from docstrings.
96.0/ best 96.0

Best in class

MATH-500Competition-style math problems.
97.8/ best 97.8

Best in class

Humanity's Last ExamHardest reasoning benchmark (with tools). ~50% = near frontier.
57.2/ best 57.9
AIME 2025American Invitational Mathematics Examination.
96.5/ best 96.5

Best in class

Other OpenAI models

O

GPT-5.5

Frontier

≈$15.00/M
O

GPT-5.1 Mini

Fast / Cheap

≈$0.70/M

Built by GLM-5.2 · Scores are directional, drawn from public mid-2026 benchmarks.

SWE-bench, GPQA, MMLU, AIME, HLE etc. are trademarks of their respective owners.