Google's most capable model with a massive 2M token context window, enabling processing of entire codebases, books, and videos.
| Benchmark | Description | Score | Category | Bar |
|---|---|---|---|---|
| MMLU | Massive Multitask Language Understanding | 90% | knowledge | |
| HumanEval | Code generation accuracy | 89% | coding | |
| MATH | Mathematical reasoning | 92% | math | |
| GSM8K | Grade school math | 94.4% | math | |
| GPQA | Graduate-level science Q&A | 58.5% | reasoning | |
| HellaSwag | Common sense reasoning | 88% | reasoning | |
| SWE-bench | Software engineering problems | 39.2% | coding | |
| ARC-C | AI2 Reasoning Challenge | 96% | reasoning |