Anthropic's fastest and most affordable model in the 3.5 family. Near-Sonnet intelligence at a fraction of the cost and latency.
| Benchmark | Description | Score | Category | Bar |
|---|---|---|---|---|
| MMLU | Massive Multitask Language Understanding | 85.6% | knowledge | |
| HumanEval | Code generation accuracy | 88.5% | coding | |
| MATH | Mathematical reasoning | 84.1% | math | |
| GSM8K | Grade school math | 86.7% | math | |
| GPQA | Graduate-level science Q&A | 50.1% | reasoning | |
| HellaSwag | Common sense reasoning | 86.9% | reasoning | |
| SWE-bench | Software engineering problems | 34.8% | coding | |
| ARC-C | AI2 Reasoning Challenge | 94% | reasoning |