AI Nexus Benchmarks

2025 Frontier

Curated by Antigravity (Gemini 3.6 Flash)

18 Models Tracked
•
9 Major AI Labs
•
Live LMSYS Arena ELO
Sort:
Filter by Lab:Showing 18 of 18 models
Min Context:
OpenAI

o3-mini

ReasoningFrontier

OpenAI latest compact reasoning model optimized for STEM, math, and coding with customizable compute budget per request.

Chatbot Arena ELO1362
Access TypeProprietary API
MMLU-Pro Reasoning84.5%
GPQA Graduate Science79.2%
HumanEval Coding92.4%
200k Context$1.1/M In95 tok/s
DeepSeek

DeepSeek R1

ReasoningFrontier

First-of-its-kind open-weights reasoning model utilizing large-scale reinforcement learning to produce long-chain thought traces.

Chatbot Arena ELO1360
Access TypeOpen Weights
MMLU-Pro Reasoning84%
GPQA Graduate Science75.7%
HumanEval Coding92.8%
128k Context$0.55/M In65 tok/s
OpenAI

OpenAI o1

Reasoning

Full-size flagship reasoning model trained with chain-of-thought reinforcement learning across text and visual inputs.

Chatbot Arena ELO1358
Access TypeProprietary API
MMLU-Pro Reasoning83.8%
GPQA Graduate Science75.7%
HumanEval Coding92.4%
200k Context$15/M In40 tok/s
Google

Gemini 2.0 Flash Thinking

ReasoningFrontier

Google DeepMind experimental reasoning model featuring transparent thoughts, 1M context, and high speed at ultra-low price.

Chatbot Arena ELO1345
Access TypeProprietary API
MMLU-Pro Reasoning81.2%
GPQA Graduate Science71.4%
HumanEval Coding93.1%
1,000k Context$0.15/M In110 tok/s
Anthropic

Claude 3.5 Sonnet

Frontier

Anthropic flagship model widely considered the leading software engineering model for agentic coding and nuanced prose.

Chatbot Arena ELO1334
Access TypeProprietary API
MMLU-Pro Reasoning78%
GPQA Graduate Science65%
HumanEval Coding93.7%
200k Context$3/M In72 tok/s
DeepSeek

DeepSeek V3

Frontier

A 671B parameter MoE open model requiring only 37B active parameters per token, delivering frontier quality at minimal operational cost.

Chatbot Arena ELO1318
Access TypeOpen Weights
MMLU-Pro Reasoning75.9%
GPQA Graduate Science59.1%
HumanEval Coding82.6%
128k Context$0.14/M In85 tok/s
Qwen

Qwen 2.5 Max

Alibaba Cloud flagship model built with massive pre-training scale, delivering top international benchmark standings.

Chatbot Arena ELO1310
Access TypeProprietary API
MMLU-Pro Reasoning77.2%
GPQA Graduate Science60.5%
HumanEval Coding90%
128k Context$0.4/M In78 tok/s
Google

Gemini 2.0 Flash

Frontier

Google next-gen workhorse model engineered for low-latency multimodal streaming, function calling, and fast web search integration.

Chatbot Arena ELO1308
Access TypeProprietary API
MMLU-Pro Reasoning74.8%
GPQA Graduate Science57.8%
HumanEval Coding90.8%
1,000k Context$0.1/M In140 tok/s
xAI

Grok 2

xAI flagship model trained on real-time internet data stream with frontier multimodal capabilities.

Chatbot Arena ELO1290
Access TypeProprietary API
MMLU-Pro Reasoning72.8%
GPQA Graduate Science56%
HumanEval Coding88.4%
128k Context$2/M In85 tok/s
OpenAI

GPT-4o

OpenAI omni model accepting any combination of text, audio, and image inputs and generating text, audio, and image outputs.

Chatbot Arena ELO1287
Access TypeProprietary API
MMLU-Pro Reasoning72.6%
GPQA Graduate Science53.6%
HumanEval Coding90.2%
128k Context$2.5/M In80 tok/s
Meta

Llama 3.1 405B

Meta flagship 405 billion parameter open-weights model designed to rival top closed frontier foundation models.

Chatbot Arena ELO1280
Access TypeOpen Weights
MMLU-Pro Reasoning73.3%
GPQA Graduate Science51.1%
HumanEval Coding89%
128k Context$1.2/M In35 tok/s
Meta

Llama 3.3 70B

Frontier

Meta latest 70B open model providing performance rivaling previous generation top-tier models while maintaining modest hardware requirements.

Chatbot Arena ELO1272
Access TypeOpen Weights
MMLU-Pro Reasoning70.1%
GPQA Graduate Science48.7%
HumanEval Coding88.4%
128k Context$0.2/M In90 tok/s
Anthropic

Claude 3.5 Haiku

Anthropic fastest model, matching Claude 3 Opus speed while outperforming it across key programming and instruction benchmarks.

Chatbot Arena ELO1264
Access TypeProprietary API
MMLU-Pro Reasoning67.5%
GPQA Graduate Science41.5%
HumanEval Coding88.1%
200k Context$0.8/M In125 tok/s
Google

Gemini 1.5 Pro

Google DeepMind breakthrough long-context model capable of processing up to 2 hours of audio or 1 hour of video in a single prompt.

Chatbot Arena ELO1260
Access TypeProprietary API
MMLU-Pro Reasoning69.1%
GPQA Graduate Science46.2%
HumanEval Coding84.1%
2,000k Context$1.25/M In60 tok/s
Mistral

Mistral Large 2

Mistral AI flagship 123B model designed specifically for high code comprehension and multilingual accuracy.

Chatbot Arena ELO1255
Access TypeAPI & Open Weights
MMLU-Pro Reasoning67.2%
GPQA Graduate Science46.8%
HumanEval Coding92%
128k Context$2/M In68 tok/s
Qwen

Qwen 2.5 Coder 32B

Open-weights coding champion designed specifically for software engineering that runs locally on modest hardware.

Chatbot Arena ELO1240
Access TypeOpen Weights
MMLU-Pro Reasoning64%
GPQA Graduate Science42%
HumanEval Coding92.7%
128k Context$0.1/M In105 tok/s
OpenAI

GPT-4o mini

OpenAI lightweight multimodal model replacing GPT-3.5 Turbo with far higher intelligence at a fraction of the price.

Chatbot Arena ELO1221
Access TypeProprietary API
MMLU-Pro Reasoning62.4%
GPQA Graduate Science40.2%
HumanEval Coding87.2%
128k Context$0.15/M In130 tok/s
Cohere

Command R+

Cohere flagship open-weights enterprise model specialized for RAG workflows with multi-step tool use.

Chatbot Arena ELO1215
Access TypeAPI & Open Weights
MMLU-Pro Reasoning61.5%
GPQA Graduate Science38%
HumanEval Coding75%
128k Context$2.5/M In60 tok/s
AI Nexus Benchmarks

Open-source intelligence & benchmark dashboard indexing LMSYS Chatbot Arena, MMLU-Pro, GPQA, HumanEval, and MATH-500 scores across leading AI research labs.

Built by Antigravity (Gemini 3.6 Flash)
Data updated July 2026 • Verified benchmark evaluation suites