Frontier BenchAI model benchmarks · Jul 2026
ModelsCompareLabsMethodology
All models
O

GPT-5.1 Mini

OpenAI

Workhorse small model — cheap, fast, multimodal, ubiquitous.

Fast / CheapProprietary200K contextReleased Oct 2025textvisionaudio-incode

API price · per 1M tokens

$0.4/$1.6

input / output

≈ $0.70/M blended

Benchmarks

Bar shows this model's score; tick shows the best score across all tracked models for context.

SWE-bench VerifiedReal GitHub issues, human-verified patches. The standard coding-agent benchmark.
70.0/ best 90.1
GPQA DiamondGoogle-proof graduate-level science Q&A. General reasoning proxy.
75.0/ best 94.4
MMLU-ProMulti-task language understanding across 120+ academic subjects.
80.0/ best 93.0
HumanEval+Functional code generation from docstrings.
90.0/ best 96.0
MATH-500Competition-style math problems.
86.0/ best 97.8

Other OpenAI models

O

GPT-5.5

Frontier

≈$15.00/M
O

GPT-5.5 Pro

Frontier

≈$26.25/M

Built by GLM-5.2 · Scores are directional, drawn from public mid-2026 benchmarks.

SWE-bench, GPQA, MMLU, AIME, HLE etc. are trademarks of their respective owners.