ModelLab
ModelsCompare
ModelsLlama 4 Maverick
∞

Llama 4 Maverick

Meta

MoE workhorse — efficient and openly deployable.

Mid-tierOpen weightsVision
Lab page

Released

Jan 15, 2026

Context

1M

Pricing

Open

per 1M tok

Best at

MoE efficiency

Performance

Benchmarks

Chatbot Arena
1417#13/17

Crowdsourced head-to-head chat preference Elo from LMArena. Captures perceived helpfulness and vibe across real-world prompts.

SWE-bench Verified
54.3#11/17

Share of real GitHub issues a model can autonomously resolve end-to-end (with tests passing). The de-facto agentic coding benchmark.

AIME 2025
84.0#13/16

American Invitational Mathematics Examination — competition math that rewards deep, multi-step reasoning. Reported as % solved.

GPQA Diamond
77.9#13/17

Google-proof graduate-level science Q&A (physics, biology, chemistry). PhD-level questions that resist web lookup.

MMLU-Pro
79.0#11/17

Harder, 10-way multiple-choice version of MMLU across 14 academic and professional domains. Breadth-of-knowledge signal.

Aider Polyglot
69.8#12/17

Editing a real codebase across multiple languages. Measures practical, instruction-following coding skill (not just generation).

OSWorld
20.0#12/16

Computer-use / agentic benchmark: completing real desktop OS tasks (apps, files, browsers). The leading autonomy metric.

Value score: 66/100

Quality-per-dollar blend of Chatbot Arena Elo against API price. Free / open models are scored against a floor.

Capabilities

  • MoE efficiency
  • Open weights
  • Great for fine-tuning
TextVisionCode

Closest rivals

By Chatbot Arena Elo

  • X

    Grok 4 Fast

    xAI

    1422
  • Q

    Qwen3-235B

    Qwen

    1408
  • M

    Mistral Large 3

    Mistral

    1429

More from

Meta AI

∞

Llama 4 Behemoth

Meta's largest open-weights model, near-frontier quality.

Data last checked May 30, 2026 (2mo ago)

ModelLab

An independent dashboard for comparing frontier AI models. Benchmark scores are drawn from public model cards, papers, and leaderboards. Always verify against the original source before making decisions.

Navigate

  • All models
  • Compare
  • About the data

Sources

  • LMArena ↗
  • SWE-bench ↗
  • LiveBench ↗

Data snapshot last checked June 17, 2026.

Not affiliated with any AI lab. Scores are approximations.