Gemini 3.1 Pro
Google DeepMindFeb 19, 2026N/A
Google's deep-reasoning upgrade: verified 77.1% on ARC-AGI-2 (more than double Gemini 3 Pro) and ~85% average accuracy at 128k-token retrieval. The strongest hard-reasoning score we track.
- Status
- Generally available
- License
- Proprietary API
- Context window
- 1M
- Max output
- 64K
- Input · $/M
- $2
- Output · $/M
- $12
- Cached input · $/M
- $0.2
- Blended · $/M
- $4.50
- Parameters
- Undisclosed
- Inputs
- Text + images, audio, video
Benchmarks
Composite 77.1· mean of published scores
Reasoning & knowledge
GPQA Diamond
no public score
MMLU-Pro
no public score
Humanity's Last Exam
no public score
ARC-AGI-2Leader
77.1%
Coding
SWE-bench Verified
no public score
SWE-bench Pro
no public score
LiveCodeBench
no public score
Agents & computer use
Terminal-Bench 2
no public score
OSWorld 2.0
no public score
BrowseComp
no public score
Math
AIME 2026
no public score
Known strengths
- ARC-AGI-2 leader (77.1%)
- Groundedness lead vs peers
- Native video understanding
- 1M context
Caveats
- Launched in preview Feb 2026; GA followed later
- Effective long-context quality can degrade well before 1M tokens