AI
ModelBench
Compare

Compare the Best AI Models

Benchmarks, pricing, and capabilities for the latest frontier AI models from OpenAI, Anthropic, Google, Meta, and more.

Explore ModelsCompare Models
16
Models Tracked
9
AI Labs
8
Benchmarks
Best Overall
GPT-4.5
Highest MMLU at 91.5%
Best Value
Gemini 2.0 Flash
Just $0.10/M input tokens
Best Coder
GPT-4o
90.2% on HumanEval
Longest Context
Gemini 1.5 Pro
2M token window
18 models
G

GPT-4.5

Transformer (MoE)

OpenAI

OpenAI's largest and most capable general-purpose model with deep knowledge and reasoning. Delivers state-of-the-art on many benchmarks.

Context:128K
Params:Unknown
Input:$75/M
Output:$150/M
textvision
Coding
92.1%
Math
94.5%
Overall
86%
Details
o

o1 Pro

Chain-of-Thought Transformer

OpenAI

Premium reasoning model for high-stakes tasks requiring the highest accuracy. Offers extended thinking with up to 100K tokens of reasoning.

Context:200K
Params:Unknown
Input:$150/M
Output:$600/M
Coding
91.8%
Math
96.4%
Overall
85%
Details
o

o3-mini

Chain-of-Thought Transformer

OpenAI

OpenAI's efficient reasoning model optimized for scientific, math, and coding tasks with chain-of-thought reasoning at lower cost.

Context:200K
Params:Unknown
Input:$1.1/M
Output:$4.4/M
Coding
93.4%
Math
94%
Overall
84%
Details
C

Claude 3.5 Sonnet

Transformer

Anthropic

Anthropic's most intelligent model in the 3.5 family, featuring Artifacts for creating and editing applications in real-time.

Context:200K
Params:~81B activated (MoE)
Input:$3/M
Output:$15/M
Coding
92%
Math
88.3%
Overall
81%
Details
G

Gemini 1.5 Pro

Transformer (MoE)

Google

Google's most capable model with a massive 2M token context window, enabling processing of entire codebases, books, and videos.

Context:2.0M
Params:Unknown
Input:$1.25/M
Output:$5/M
textvisionaudiovideo
Coding
89%
Math
92%
Overall
81%
Details
G

Gemini 2.0 Flash

Transformer (MoE)

Google

Google's fastest and most capable model with native multimodality, extreme context window, and built-in function calling for agents.

Context:1.0M
Params:Unknown
Input:$0.1/M
Output:$0.4/M
textvisionaudiofunction-calling
Coding
89.7%
Math
90%
Overall
80%
Details
G

GPT-4o

Transformer (MoE)

OpenAI

OpenAI's flagship multimodal model capable of processing text, vision, and audio inputs. Excels at real-time voice conversations and multimodal reasoning.

Context:128K
Params:Unknown
Input:$2.5/M
Output:$10/M
textvisionaudio
Coding
90.2%
Math
90%
Overall
79%
Details
L

Llama 3.1 405B

Transformer (MoE)

Meta

Meta's largest open-weight model, rivaling GPT-4 class performance with full open weights for commercial and research use.

Context:128K
Params:405B
Input:$0/M
Output:$0/M
Coding
89%
Math
90%
Overall
79%
Details
C

Claude 3 Opus

Transformer

Anthropic

Anthropic's most powerful model at launch, with expert-level performance across math, coding, and reasoning.

Context:200K
Params:~81B activated (MoE)
Input:$15/M
Output:$75/M
textvision
Coding
84.9%
Math
87.3%
Overall
78%
Details
G

Gemini Ultra 1.0

Transformer (MoE)

Google

Google's original most capable model, designed to be a general-purpose AI that outperforms experts on 30 of 32 benchmarks.

Context:128K
Params:Unknown
Input:$0/M
Output:$0/M
textvisionaudio
Coding
74.4%
Math
94.4%
Overall
78%
Details
L

Llama 3.3 70B

Transformer (MoE)

Meta

Meta's efficient 70B instruction model that matches or exceeds Llama 3.1 70B across key benchmarks with improved inference efficiency.

Context:128K
Params:70B
Input:$0/M
Output:$0/M
Coding
88.4%
Math
87.8%
Overall
77%
Details
D

DeepSeek-V2.5

Transformer (MoE)

DeepSeek

DeepSeek's merged model combining deep reasoning with general instruction following at extremely low cost.

Context:128K
Params:236B (21B active)
Input:$0.14/M
Output:$0.28/M
Coding
89.2%
Math
86%
Overall
77%
Details
C

Claude 3.5 Haiku

Transformer

Anthropic

Anthropic's fastest and most affordable model in the 3.5 family. Near-Sonnet intelligence at a fraction of the cost and latency.

Context:200K
Params:Unknown
Input:$1/M
Output:$5/M
Coding
88.5%
Math
84.1%
Overall
76%
Details
M

Mistral Large

Transformer (MoE)

Mistral

Mistral AI's most powerful model with excellent multilingual capabilities, strong coding, and competitive European AI development.

Context:128K
Params:129B (12B active)
Input:$2/M
Output:$6/M
textvision
Coding
87.7%
Math
83.5%
Overall
76%
Details
G

Grok-2

Transformer (MoE)

xAI

xAI's Grok-2 model with real-time knowledge access via X (Twitter) integration, featuring image generation capabilities.

Context:128K
Params:Unknown
Input:$15/M
Output:$15/M
textvision
Coding
87.6%
Math
86.3%
Overall
76%
Details
Q

Qwen 2.5 72B

Transformer

Alibaba

Alibaba's powerful open-source model with strong multilingual performance, particularly for Chinese and other Asian languages.

Context:128K
Params:72B
Input:$0/M
Output:$0/M
Coding
85.2%
Math
83.4%
Overall
74%
Details
C

Command R+

Transformer (MoE)

Cohere

Cohere's model optimized for enterprise RAG workflows with built-in search and retrieval capabilities for business applications.

Context:128K
Params:111B (27B active)
Input:$3/M
Output:$15/M
textvisionsearch
Coding
82.4%
Math
77.2%
Overall
71%
Details
P

Phi-3.5 Mini

Transformer (Dense)

Microsoft

Microsoft's highly efficient small model that punches above its weight class, demonstrating the power of quality synthetic data training.

Context:128K
Params:3.8B
Input:$0/M
Output:$0/M
Coding
79%
Math
75.6%
Overall
66%
Details
AI
ModelBench

The open benchmark platform for comparing the latest AI models and their capabilities.

Platform

  • Compare Models
  • All Models
  • Methodology

AI Labs

  • OpenAI
  • Anthropic
  • Google DeepMind
  • Meta AI

Resources

  • Documentation
  • Benchmark Guide
  • API Reference
Benchmarks are sourced from official model cards and independent evaluations. Data is for informational purposes only.
© 2026 ModelBench. Built for the AI community.