ModelLab
ModelsCompare
ModelsGemini 3.1 Pro
G

Gemini 3.1 Pro

Google

Leads on reasoning with a native 2M-token context.

FrontierVisionAudio inVoice out
Lab page

Released

May 20, 2026

Context

2M

Pricing

$2.50/$10

per 1M tok

Best at

Long-context king

Performance

Benchmarks

Chatbot Arena
1505#3/17

Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.

SWE-bench Verified
72.0#5/17

Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.

AIME 2025Leader
97.1#1/16

American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.

GPQA DiamondLeader
91.2#1/17

Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.

MMLU-Pro
86.4#3/17

Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.

Aider Polyglot
80.3#5/17

Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).

OSWorld
40.6#3/16

Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.

Value score: 77/100

Quality-per-dollar blend of Chatbot Arena Elo against API price. Free / open models are scored against a floor.

API pricing

Input$2.50per 1M tokens
Output$10per 1M tokens
1M tokens blended$13in + out

Capabilities

  • Long-context king
  • Top reasoning
  • Multimodal native
TextVisionAudio inputVoice outputReasoningCode

Closest rivals

By Chatbot Arena Elo

  • OA

    GPT-5.5

    OpenAI

    1512
  • X

    Grok 4.3

    xAI

    1498
  • A

    Claude Opus 4.8

    Anthropic

    1521

More from

Google DeepMind

G

Gemini 3 Flash

Best price-to-performance in the lineup, huge context.

Data last checked Jun 14, 2026 (1mo ago)

ModelLab

An independent dashboard for comparing frontier AI models. Benchmark scores are drawn from public model cards, papers, and leaderboards. Always verify against the original source before making decisions.

Navigate

  • All models
  • Compare
  • About the data

Sources

  • LMArena ↗
  • SWE-bench ↗
  • LiveBench ↗

Data snapshot last checked June 17, 2026.

Not affiliated with any AI lab. Scores are approximations.