Claude 3.7 Sonnet
APIHybrid reasoning frontier model combining instant response with granular extended thinking
Hybrid reasoning frontier standard
Solves 7 out of 10 real GitHub issues
1M context + 135 tok/sec throughput
MIT Licensed reasoning breakthrough
Hybrid reasoning frontier model combining instant response with granular extended thinking
Flagship reasoning model optimized for deep science, mathematics, and complex multi-step logic
Googleβs frontier model combining 2-million context window with complex reasoning and multimodal analysis
High-speed reasoning model tailored for STEM, coding, and cost-effective chain-of-thought
Open-weights reasoning breakthrough trained with pure large-scale RL, rivaling proprietary titans
Alibabaβs frontier flagship challenging the worldβs best models in English, Chinese, coding, and math
Real-time ultra-fast multimodal model with 1M context, native audio/vision streaming, and extreme cost efficiency
Benchmark standard for coding, agentic autonomy, and human-like natural prose
Versatile omnimodal workhorse model powering ChatGPT with native vision, audio, and text
xAIβs flagship model with real-time X integration, sharp reasoning, and visual analysis
671B parameter Mixture-of-Experts open-weights base model redefining LLM cost and architecture
The open code specialist that matches or surpasses closed models in real-world programming benchmarks
Metaβs best open model per compute, delivering 405B-level capabilities at 70B footprint
European flagship with 123B parameters, advanced multilingual proficiency, and enterprise governance
Economical small model supporting multimodal inputs with rapid execution
Blazing speed and remarkable intelligence at a fraction of the frontier cost
Crowdsourced blind A/B evaluation by LMSYS with over 2M human comparisons.
Human-validated subset of SWE-bench resolving real-world GitHub issues across large Python repos.
Graduate-level Google-proof Q&A benchmark curated by biology, physics, and chemistry PhDs.
Harder, 14-subject multi-choice reasoning benchmark with 10 options per question (drastically reducing guessing luck).
Challenging 500-problem subset of the Hendrycks MATH benchmark covering high school competition level math.
American Invitational Mathematics Examination competition problems requiring advanced multi-step proof search.
Massive Multi-discipline Multimodal Understanding benchmark spanning 30 college-level subjects requiring image comprehension.