ApertureAug 2026
FieldCompareLabsGuideMethod
  • Field
  • Compare
  • Labs
  • Guide
Anthropic
A

Claude Sonnet 5

Jun 2026 · Proprietary · Current · supported evidence

Anthropic’s daily driver. Not the crown, but the one you actually leave on in production if Opus is too rich.

Open compareLab site

Intelligence

55.3

AA Index

API price

$2 / $10

Input / output per 1M

Context

1M

82 tok/s

Capability

Intelligence Index55.3
Coding index71.5
Agentic index49.7
Arena Elo (offset)158

Use it when

  • ▸Product chat
  • ▸Docs
  • ▸Everyday coding

Skip if

  • –You are already paying for Opus 5 and need the extra points

Benchmarks

GPQA Diamond
Graduate-level science questions designed so Google search is not enough. Still one of the cleanest knowledge/reasoning splits.
—
Humanity’s Last Exam
Expert-written questions across fields. Harder than MMLU; the current differentiator for “does this model actually know things.”
—
MMLU-Pro
Harder, less-saturated successor to MMLU. Classic MMLU is above 90% for every flagship and no longer ranks the field.
—
SWE-bench Verified
500 human-validated GitHub issues. Score swings 5–15 points by harness — treat vendor numbers as an upper bound.
85.2%
SWE-bench Pro
Harder, contamination-resistant coding eval. Currently the best public split between “can code” and “can maintain a repo.”
63.2%
Terminal-Bench 2.1
End-to-end tasks in a real terminal. Better proxy for coding agents than HumanEval, which is fully saturated.
80.4%
ARC-AGI-2
Abstract visual puzzles. Rewards generalization over memorization. GPT-5.6 Sol currently leads the published set.
—
AIME 2025
American Invitational Mathematics Examination. Contest math; reasoning-mode models dominate.
—
MMMU
College-level multimodal understanding across diagrams, charts, and exam figures.
—

Modalities

text · vision · tools

Reasoning mode

hybrid

Hybrid and reasoning models spend tokens thinking. That raises GPQA and agents, and also raises latency and bill.

Close on intelligence

  • OGPT-5.6 Terra56.6
  • OGPT-5.6 Luna52.3
  • GGemini 3.7 Flash56

Same lab

  • Claude Mythos 5
  • Claude Fable 5
  • Claude Opus 5
  • Claude Opus 4.8

Aperture is a field guide, not a vendor. Scores compiled 28 Aug 2026.

MethodologyHow to pick