On May 1, 2026, the Center for AI Standards and Innovation (CAISI), housed at the National Institute of Standards and Technology (NIST), published its official evaluation of DeepSeek V4 Pro, the open-weight Chinese model. The one-line conclusion: it is the most capable PRC model CAISI has evaluated to date, yet it lags the US frontier by about eight months, landing at roughly the level of OpenAI’s GPT-5.
What makes the report worth a developer’s attention is not just the ranking but the methodology. CAISI uses pre-committed benchmark suites and Item Response Theory (IRT) weighting to squeeze out contamination and cherry-picking. In a year where self-reported scores are increasingly incomparable, a government-grade, auditable evaluation procedure is taking shape — and it is on track to plug directly into pre-release review processes.
How the Evaluation Works
Three design choices stand out. First, benchmarks are pre-committed: the suite is fixed before any runs, not selected after seeing results. Second, CAISI uses held-out, non-public benchmarks — including its internally built PortBench and ARC-AGI-2’s semi-private set — to reduce training-data contamination risk. Third, it controls token budgets and scaffolding so models answer under comparable conditions, then estimates aggregate capability with an IRT-inspired method. Five domains are covered: cybersecurity, software engineering, natural sciences, abstract reasoning, and mathematics.
Headline Results: Elo 800 and a Five-Domain Profile
The core number: DeepSeek V4 Pro posts an IRT-estimated Elo of 800 (±28), against 1260 for GPT-5.5, 999 for Opus 4.6, and 749 for GPT-5.4 mini. The domain profile is lopsided. In math it nearly matches top US models: 97 percent on OTIS-AIME-2025, and 96 percent on both PUMaC 2024 and SMT 2025. Abstract reasoning is another story — 46 percent on the ARC-AGI-2 semi-private set — and cybersecurity manages only 32 percent on CTF-Archive-Diamond (imputed via IRT). Software engineering and the natural sciences also trail the frontier. The three weakest areas are, inconveniently, three of the ones agentic workloads depend on most.
The Gap With DeepSeek’s Own Claims
The other story is the gap between vendor narrative and independent measurement. DeepSeek’s own technical report pitched V4 Pro as on par with Opus 4.6 and GPT-5.4 — models released roughly two months before the evaluation. CAISI’s testing places it back at eight-month-old GPT-5 territory. As The Decoder’s Matthias Bastian notes, this is not necessarily proof of dishonesty; it is the systematic distance between a marketing document with hand-picked benchmarks and comparators, and a pre-registered evaluation. Procurement decisions should weight the latter.
Cost Edge and Political Context
Cost is where DeepSeek genuinely wins. Against the only US reference model meeting CAISI’s filtering criteria, GPT-5.4 mini, DeepSeek V4 was cheaper on five of seven benchmarks, ranging from 53 percent less expensive to 41 percent more expensive. With agent workloads burning millions of tokens and premium US models getting pricier, that spread reorders shortlists — Cursor has already fine-tuned a cheaper coding model on the Chinese open-weight Kimi K2.5. Sam Altman himself acknowledged the tension on X, writing that “just being smarter is still the most important thing” — a reminder that capability-per-dollar cuts both ways when the capability gap is real. But the political framing deserves skepticism: The Decoder points out that Artificial Analysis’s Intelligence Index shows the US-China gap has held roughly constant, and that CAISI, as a government agency, “likely has its own political agenda.” The open-source route divergence flagged in our opening-of-year outlook now comes with an official-evaluation layer on top.
What It Means for Developers
Three practical takeaways. First, select by domain, not by total score: math-heavy pipelines can seriously consider open-weight models; security-sensitive and long-chain reasoning tasks should stay conservative for now. Second, borrow CAISI’s methodology — pre-committed benchmarks, held-out items, fixed token budgets — because that procedure resists contamination far better than another lap of public leaderboards. Third, note the timing: the same week, CAISI announced expanded pre-release evaluation of US frontier models. Government benchmarking is graduating from research tool to regulatory infrastructure, and understanding how it scores models is understanding tomorrow’s compliance bar.
Sources
- CAISI Evaluation of DeepSeek V4 Pro — NIST
- China is falling behind in the AI race, according to a US government benchmark — The Decoder
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
