DeepSeek-V4-Pro (0813)
DeepSeekAug 13, 2026#1
NewOpen weights
MIT-licensed frontier open weights refreshed August 13. Official report: 80.6% SWE-bench Verified, 90.1% GPQA Diamond, 87.5 MMLU-Pro, 93.5 LiveCodeBench — astonishing numbers at an order of magnitude under US frontier pricing. Independent harnesses (e.g. DeepSWE) show softer results.
- Status
- Generally available
- License
- Open weights · MIT
- Context window
- 1M
- Max output
- 64K
- Input · $/M
- $0.7
- Output · $/M
- $3.48
- Cached input · $/M
- $0.07
- Blended · $/M
- $1.40
- Parameters
- 1.6T-class MoE (sources cite 1.0–1.6T)
- Inputs
- Text only
Benchmarks
Composite 87.9· mean of published scores
Reasoning & knowledge
GPQA Diamond
90.1%
dataset best 93.6MMLU-ProLeader
87.5%
Humanity's Last Exam
no public score
ARC-AGI-2
no public score
Coding
SWE-bench Verified
80.6%
dataset best 93.4SWE-bench Pro
no public score
LiveCodeBenchLeader
93.5%
Agents & computer use
Terminal-Bench 2
no public score
OSWorld 2.0
no public score
BrowseComp
no public score
Math
AIME 2026
no public score
Known strengths
- Cheapest per-point frontier-class coder
- Open MIT weights you can host anywhere
- Full 1M context without surcharge
Caveats
- Headline SWE-V is self-reported; stricter third-party harnesses score lower
- Parameter count reporting varies between 1.0T and 1.6T