← Back to explorer
GoogleReleased 2026-03

Gemini 3.1 Pro

Google DeepMind science and long-context leader. Highest GPQA Diamond in this snapshot with a 2M-token context window.

Context
2M
Speed
85 t/s
Price /M
$2.5 / $15
Access
api

When to use it

Scientific reasoning, huge document corpora, and Google Cloud / Workspace stacks.

Watch-outs

Coding Arena trails Claude/OpenAI leaders; verify agent harness fit.

ScienceLong contextMultimodalAgents

Benchmarks

Arena Elo1492

Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.

AA Intelligence Index54

Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).

SWE-bench Verified80.6%

Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.

GPQA Diamond94.3%

PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.

MMLU-Pro92.6%

Harder multi-choice knowledge/reasoning suite than classic MMLU.

Humanity's Last Exam44.7%

Frontier closed-ended academic difficulty across many domains.

More from Google

  • Gemini 3 Flash

    High-throughput apps, classification, and cost-sensitive Google stack workloads.

    AA 42
  • Gemini 3.0

    Balanced Google Cloud apps needing solid quality at lower spend.

    AA 45

Model Gauge · built by Cursor Grok 4.5 High · curated snapshot, not a live leaderboard

Scores change weekly — verify critical decisions against primary sources.