Back to all models
o

o3-mini

OpenAIReleased 2025-01-31
Official Page

OpenAI's efficient reasoning model optimized for scientific, math, and coding tasks with chain-of-thought reasoning at lower cost.

Overall Score
84%
Cost Efficiency
100
Context Window
200K

Strengths

  • ✓Exceptional reasoning capabilities
  • ✓Low cost for a reasoning model
  • ✓Strong math and coding

Weaknesses

  • ✗Text-only
  • ✗Not as versatile as general models
Architecture
Chain-of-Thought Transformer
Parameters
Unknown
Input
$1.1/M
Output
$4.4/M

Benchmark Performance

BenchmarkDescriptionScoreCategoryBar
MMLUMassive Multitask Language Understanding85.9%knowledge
HumanEvalCode generation accuracy93.4%coding
MATHMathematical reasoning94%math
GSM8KGrade school math96%math
GPQAGraduate-level science Q&A70.7%reasoning
HellaSwagCommon sense reasoning89.9%reasoning
SWE-benchSoftware engineering problems45.1%coding
ARC-CAI2 Reasoning Challenge97%reasoning
AI
ModelBench

The open benchmark platform for comparing the latest AI models and their capabilities.

Platform

  • Compare Models
  • All Models
  • Methodology

AI Labs

  • OpenAI
  • Anthropic
  • Google DeepMind
  • Meta AI

Resources

  • Documentation
  • Benchmark Guide
  • API Reference
Benchmarks are sourced from official model cards and independent evaluations. Data is for informational purposes only.
© 2026 ModelBench. Built for the AI community.