DeepSeek V4 Pro
Open-weights value king. Near-frontier intelligence at a fraction of closed-lab price — standout on AA cost charts.
When to use it
Self-hosting, cost-sensitive production, and research fine-tunes.
Watch-outs
Ops burden for hosting; safety/tooling polish varies by deployment.
Benchmarks
Human preference ranking from blind pairwise chats. Higher is better; top frontier models cluster within ~50–80 Elo.
Artificial Analysis composite across agents, coding, science, and general evaluations (v4.1 weighting).
Percent of real GitHub issues resolved end-to-end. Strong signal for agentic coding usefulness.
PhD-level science questions. Separates frontier reasoning models better than saturated knowledge tests.
Harder multi-choice knowledge/reasoning suite than classic MMLU.
Frontier closed-ended academic difficulty across many domains.