Back to all models
o

o1 Pro

OpenAIReleased 2024-12-05
Official Page

Premium reasoning model for high-stakes tasks requiring the highest accuracy. Offers extended thinking with up to 100K tokens of reasoning.

Overall Score
85%
Cost Efficiency
13
Context Window
200K

Strengths

  • ✓Highest accuracy reasoning
  • ✓Extensive chain-of-thought
  • ✓Best for math and science

Weaknesses

  • ✗Very expensive
  • ✗Slow inference
Architecture
Chain-of-Thought Transformer
Parameters
Unknown
Input
$150/M
Output
$600/M

Benchmark Performance

BenchmarkDescriptionScoreCategoryBar
MMLUMassive Multitask Language Understanding87.3%knowledge
HumanEvalCode generation accuracy91.8%coding
MATHMathematical reasoning96.4%math
GSM8KGrade school math96.7%math
GPQAGraduate-level science Q&A77.6%reasoning
HellaSwagCommon sense reasoning88.6%reasoning
SWE-benchSoftware engineering problems44%coding
ARC-CAI2 Reasoning Challenge97%reasoning
AI
ModelBench

The open benchmark platform for comparing the latest AI models and their capabilities.

Platform

  • Compare Models
  • All Models
  • Methodology

AI Labs

  • OpenAI
  • Anthropic
  • Google DeepMind
  • Meta AI

Resources

  • Documentation
  • Benchmark Guide
  • API Reference
Benchmarks are sourced from official model cards and independent evaluations. Data is for informational purposes only.
© 2026 ModelBench. Built for the AI community.