Gemini 3.1 Pro
Google DeepMind science and long-context leader. Highest GPQA Diamond in this snapshot with a 2M-token context window.
When to use it
Scientific reasoning, huge document corpora, and Google Cloud / Workspace stacks.
Watch-outs
Coding Arena trails Claude/OpenAI leaders; verify agent harness fit.
Benchmarks
Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.
Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).
Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.
PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.
Harder multi-choice knowledge/reasoning suite than classic MMLU.
Frontier closed-ended academic difficulty across many domains.
More from Google
- Gemini 3 FlashAA 42
High-throughput apps, classification, and cost-sensitive Google stack workloads.
- Gemini 3.0AA 45
Balanced Google Cloud apps needing solid quality at lower spend.