ModelLab
ModelsCompare
ModelsClaude Opus 4.8
A

Claude Opus 4.8

Anthropic

Top of the LLM Stats overall board; leading on deep coding.

FrontierVision
Lab page

Released

Jun 3, 2026

Context

500K

Pricing

$15/$75

per 1M tok

Best at

#1 overall

Performance

Benchmarks

Chatbot ArenaLeader
1521#1/17

Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.

SWE-bench VerifiedLeader
81.2#1/17

Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.

AIME 2025
95.8#3/16

American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.

GPQA Diamond
90.4#2/17

Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.

MMLU-ProLeader
87.9#1/17

Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.

Aider PolyglotLeader
88.5#1/17

Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).

OSWorldLeader
44.1#1/16

Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.

Value score: 67/100

Quality-per-dollar blend of Chatbot Arena Elo against API price. Free / open models are scored against a floor.

API pricing

Input$15per 1M tokens
Output$75per 1M tokens
1M tokens blended$90in + out

Capabilities

  • #1 overall
  • Best for coding
  • 500K context
TextVisionReasoningCode

Closest rivals

By Chatbot Arena Elo

  • OA

    GPT-5.5

    OpenAI

    1512
  • G

    Gemini 3.1 Pro

    Google

    1505
  • X

    Grok 4.3

    xAI

    1498

More from

Anthropic

A

Claude Sonnet 4.6

The model most developers actually ship on.

A

Claude Haiku 4.6

Sub-second responses at near-frontier quality.

Data last checked Jun 17, 2026 (1mo ago)

ModelLab

An independent dashboard for comparing frontier AI models. Benchmark scores are drawn from public model cards, papers, and leaderboards. Always verify against the original source before making decisions.

Navigate

  • All models
  • Compare
  • About the data

Sources

  • LMArena ↗
  • SWE-bench ↗
  • LiveBench ↗

Data snapshot last checked June 17, 2026.

Not affiliated with any AI lab. Scores are approximations.