Side-by-side benchmarks for GPT, Claude, Gemini, Llama, Mistral, Grok & more. Start from your use case, not the hype — filter by modality, price, and the benchmarks that matter to you.
MMLU / MMLU-Pro for breadth, GPQA Diamond for grad-level science, MATH + AIME for elite math, HumanEval + SWE-Bench Verified for real code, MMMU for vision, IFEval for instruction following.
Reasoning & research: sort by AIME + GPQA. Shipping features: SWE-Bench + latency + price. Long docs/video: filter 200K+ and Gemini/GPT-4o class. Self-host: open-weights + price.
Scores are noisy, prompts differ by lab, and contamination happens. Pair benchmarks with Arena ELO (human preference) and your own evals. Prices and context windows change frequently — verify before prod.