Board/DeepSeek/DeepSeek-V4-Pro (0813)
D

DeepSeek-V4-Pro (0813)

DeepSeek·Aug 13, 2026#1

NewOpen weights

MIT-licensed frontier open weights refreshed August 13. Official report: 80.6% SWE-bench Verified, 90.1% GPQA Diamond, 87.5 MMLU-Pro, 93.5 LiveCodeBench — astonishing numbers at an order of magnitude under US frontier pricing. Independent harnesses (e.g. DeepSWE) show softer results.

Status
Generally available
License
Open weights · MIT
Context window
1M
Max output
64K
Input · $/M
$0.7
Output · $/M
$3.48
Cached input · $/M
$0.07
Blended · $/M
$1.40
Parameters
1.6T-class MoE (sources cite 1.0–1.6T)
Inputs
Text only

Benchmarks

Composite 87.9· mean of published scores

Reasoning & knowledge

  • GPQA Diamond

    90.1%

    dataset best 93.6
  • MMLU-ProLeader

    87.5%

  • Humanity's Last Exam

    no public score

  • ARC-AGI-2

    no public score

Coding

  • SWE-bench Verified

    80.6%

    dataset best 93.4
  • SWE-bench Pro

    no public score

  • LiveCodeBenchLeader

    93.5%

Agents & computer use

  • Terminal-Bench 2

    no public score

  • OSWorld 2.0

    no public score

  • BrowseComp

    no public score

Math

  • AIME 2026

    no public score

Known strengths

  • Cheapest per-point frontier-class coder
  • Open MIT weights you can host anywhere
  • Full 1M context without surcharge

Caveats

  • Headline SWE-V is self-reported; stricter third-party harnesses score lower
  • Parameter count reporting varies between 1.0T and 1.6T

Primary sources

  • Spec summary (Morph)
  • Benchmark roundup (Macaron)
Compare this model →

More from DeepSeek

  • DDeepSeek-V4 FlashApr 24, 2026 · 512K ctx