AI Briefing
KO

DeepSeek V4, China's Top Model

·2026.05.03 12:10

Key point

CAISI evaluated DeepSeek V4 and found it to be China's top model, but about 8 months behind the US frontier.

1 / 2

Details

CAISI evaluated the open-weight model DeepSeek V4 Pro in April 2026 and estimated it to be about 8 months behind the US frontier.

The evaluation focused on 5 domains and 9 benchmarks, and the IRT-based comparison used 16 benchmarks and 35 models. DeepSeek V4 was classified as the best-performing model from the PRC that CAISI has evaluated so far.

  • Its IRT-estimated Elo was 800 ±28, lower than GPT-5.5 (1260 ±28) and Anthropic Opus 4.6 (999 ±27).
  • On the developer's own benchmarks, it appeared comparable to Opus 4.6 and GPT-5.4, but on CAISI's non-public benchmarks it was closer to GPT-5 level.
  • It notably underperformed compared to US models on ARC-AGI-2 semi-private, PortBench, and CTF-Archive-Diamond in particular.
  • In terms of cost, it was cheaper on 5 of 7 benchmarks compared to GPT-5.4 mini, ranging from 53% cheaper to 41% more expensive.
  • CAISI served the model using H200/B200 GPUs and the developer's recommended settings, and reproduced the developer's self-reported results on GPQA-Diamond.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.