LIVE
News

DeepSeek V4 Pro 0813 Performance Analysis: Benchmarking the Latest MoE Architecture

According to Artificial Analysis, DeepSeek V4 Pro 0813 (max) registers 53 on the Intelligence Index, a composite aggregating ten benchmarks spanning reasoning, coding, agentic tool use, and knowledge.

Shane Barrett·updated August 18, 2026

DeepSeek V4 Pro 0813 Performance Analysis: Benchmarking the Latest MoE Architecture

The score places the model one point above V4 Flash 0731, DeepSeek's smaller sibling that had briefly closed the gap. The delta to the closed-model frontier — ten points behind Claude Opus 5 (max) at 63 — remains the binding constraint.

Architecture and Frontier Position

V4 Pro debuted in April at 52, alongside V4 Flash. The 0813 build retains the 1.6 trillion parameter Mixture-of-Experts backbone with 49 billion active parameters, more than double V3's footprint. At 53, the model occupies the mid-table cluster: tied with Z.AI's GLM-5.2 (max), one point ahead of GPT-5.6 Luna (max) and Gemini 3.6 Flash at 52, fifteen points above Nemotron 3 Ultra at 38. The leading closed systems — Claude Opus 5 (63), Claude Fable 5 (62), GPT-5.6 Sol (max) and Grok 4.6 (high) tied at 61, Kimi K3 (max) at 60, Muse Spark 1.2 (xhigh) at 57 — establish a gap the index measures without adjustment for deployment cost.

Per-Benchmark Decomposition

Sub-benchmark results expose uneven capability. On GPQA Diamond, V4 Pro reaches 93%, tying Claude Opus 5 and Kimi K3 (low), trailing Grok 4.6 (high) at 95% by two points. GDPval-AA v2 yields 55%, ahead of GPT-5.6 Luna (53%) and Gemini 3.6 Flash (50%), behind Claude Opus 5 (67%) and Grok 4.6 (62%). Terminal-Bench v2.1 records 79%, a ten-point deficit against Claude Opus 5's 89%. SciCode produces 49%, the lowest score among highlighted comparators, trailing Claude Fable 5's 60% and even Nemotron 3 Ultra's 40%. Agentic tool use on τ³-Banking lands at 40%, ahead of Claude Fable 5 (39%) but behind Grok 4.6 (51%). Humanity's Last Exam registers 39%, tied with MiniMax-M3 per the evaluation. CritPt, a physics reasoning benchmark, returns 18%, well behind GPT-5.6 Sol (32%) and Claude Opus 5 (29%).

Practitioner Implications

The composite masks domain-specific trade-offs. GPQA Diamond strength does not transfer to SciCode or CritPt, where V4 Pro trails comparators ranked lower on the aggregate index. Practitioners should reproduce relevant sub-benchmarks against their workload before allocating inference budget. Cost positioning, rather than absolute capability, appears to constitute the model's intended deployment rationale; pricing data was not included in the evaluation. Independent evaluators such as Artificial Analysis operate within a broader methodological posture: verification conducted outside vendor-supplied documentation, akin to the independent infrastructure built by organizers working beyond mainstream platforms, determines practical utility more reliably than headline composite scores.