← Back to explorer
AnthropicReleased 2026-05

Claude Opus 4.8

Production-available Anthropic flagship. Tops or near-tops Arena overall and coding while remaining the most practical 'max intelligence' Claude for most teams.

Context
1M
Speed
72 t/s
Price /M
$5 / $25
Access
api

When to use it

Coding agents, computer use, and careful long-form reasoning in production.

Watch-outs

Still expensive at max effort; slower than Sonnet for high-volume traffic.

CodingAgentsWritingScience

Benchmarks

Arena Elo1512

Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.

AA Intelligence Index56

Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).

SWE-bench Verified88.6%

Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.

GPQA Diamond93.6%

PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.

MMLU-Pro

Harder multi-choice knowledge/reasoning suite than classic MMLU.

Humanity's Last Exam49.8%

Frontier closed-ended academic difficulty across many domains.

More from Anthropic

  • Claude Fable 5

    Hard coding agents, long computer-use sessions, and high-stakes knowledge work.

    AA 60
  • Claude Opus 4.7

    Complex coding agents and careful editorial writing.

    AA 52
  • Claude Sonnet 4

    Day-to-day coding copilots, customer agents, and high-volume Claude deployments.

    AA 48

Model Gauge · built by Cursor Grok 4.5 High · curated snapshot, not a live leaderboard

Scores change weekly — verify critical decisions against primary sources.