ModelLab
ModelsCompare
ModelsGPT-5.1 mini
OA

GPT-5.1 mini

OpenAI

Fast, cheap, capable — the everyday workhorse.

Small / fastVision
Lab page

Released

Nov 4, 2025

Context

200K

Pricing

$0.40/$1.60

per 1M tok

Best at

Low latency

Performance

Benchmarks

Chatbot Arena
1402#15/17

Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.

SWE-bench Verified
51.2#13/17

Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.

AIME 2025
88.5#11/16

American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.

GPQA Diamond
78.3#11/17

Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.

MMLU-Pro
78.9#12/17

Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.

Aider Polyglot
68.4#13/17

Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).

OSWorld
22.0#10/16

Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.

Value score: 56/100

Quality-per-dollar blend of Chatbot Arena Elo against API price. Free / open models are scored against a floor.

API pricing

Input$0.40per 1M tokens
Output$1.60per 1M tokens
1M tokens blended$2.00in + out

Capabilities

  • Low latency
  • Cheapest tier
  • Great for scale
TextVisionReasoning

Closest rivals

By Chatbot Arena Elo

  • Q

    Qwen3-235B

    Qwen

    1408
  • A

    Claude Haiku 4.6

    Anthropic

    1395
  • ∞

    Llama 4 Maverick

    Meta

    1417

More from

OpenAI

OA

GPT-5.5

OpenAI's flagship all-rounder with strong computer use.

OA

GPT-5.1

Previous flagship — now value-priced, still excellent.

Data last checked May 28, 2026 (2mo ago)

ModelLab

An independent dashboard for comparing frontier AI models. Benchmark scores are drawn from public model cards, papers, and leaderboards. Always verify against the original source before making decisions.

Navigate

  • All models
  • Compare
  • About the data

Sources

  • LMArena ↗
  • SWE-bench ↗
  • LiveBench ↗

Data snapshot last checked June 17, 2026.

Not affiliated with any AI lab. Scores are approximations.