DeepSeek V4 Pro 0813 Performance Analysis: Benchmarking the Latest MoE Architecture
According to Artificial Analysis, DeepSeek V4 Pro 0813 (max) registers 53 on the Intelligence Index, a composite aggregating ten benchmarks spanning reasoning, coding, agentic tool use, and knowledge.
Shane Barrett·updated August 18, 2026

The score places the model one point above V4 Flash 0731, DeepSeek's smaller sibling that had briefly closed the gap. The delta to the closed-model frontier — ten points behind Claude Opus 5 (max) at 63 — remains the binding constraint.
Architecture and Frontier Position
V4 Pro debuted in April at 52, alongside V4 Flash. The 0813 build retains the 1.6 trillion parameter Mixture-of-Experts backbone with 49 billion active parameters, more than double V3's footprint. At 53, the model occupies the mid-table cluster: tied with Z.AI's GLM-5.2 (max), one point ahead of GPT-5.6 Luna (max) and Gemini 3.6 Flash at 52, fifteen points above Nemotron 3 Ultra at 38. The leading closed systems — Claude Opus 5 (63), Claude Fable 5 (62), GPT-5.6 Sol (max) and Grok 4.6 (high) tied at 61, Kimi K3 (max) at 60, Muse Spark 1.2 (xhigh) at 57 — establish a gap the index measures without adjustment for deployment cost.
Per-Benchmark Decomposition
Sub-benchmark results expose uneven capability. On GPQA Diamond, V4 Pro reaches 93%, tying Claude Opus 5 and Kimi K3 (low), trailing Grok 4.6 (high) at 95% by two points. GDPval-AA v2 yields 55%, ahead of GPT-5.6 Luna (53%) and Gemini 3.6 Flash (50%), behind Claude Opus 5 (67%) and Grok 4.6 (62%). Terminal-Bench v2.1 records 79%, a ten-point deficit against Claude Opus 5's 89%. SciCode produces 49%, the lowest score among highlighted comparators, trailing Claude Fable 5's 60% and even Nemotron 3 Ultra's 40%. Agentic tool use on τ³-Banking lands at 40%, ahead of Claude Fable 5 (39%) but behind Grok 4.6 (51%). Humanity's Last Exam registers 39%, tied with MiniMax-M3 per the evaluation. CritPt, a physics reasoning benchmark, returns 18%, well behind GPT-5.6 Sol (32%) and Claude Opus 5 (29%).
Practitioner Implications
The composite masks domain-specific trade-offs. GPQA Diamond strength does not transfer to SciCode or CritPt, where V4 Pro trails comparators ranked lower on the aggregate index. Practitioners should reproduce relevant sub-benchmarks against their workload before allocating inference budget. Cost positioning, rather than absolute capability, appears to constitute the model's intended deployment rationale; pricing data was not included in the evaluation. Independent evaluators such as Artificial Analysis operate within a broader methodological posture: verification conducted outside vendor-supplied documentation, akin to the independent infrastructure built by organizers working beyond mainstream platforms, determines practical utility more reliably than headline composite scores.