US government benchmark shows China AI falling behind
Key point
In CAISI's evaluation, DeepSeek V4 Pro lagged about 8 months behind the leading US model.
Details
CAISI evaluated DeepSeek V4 Pro and found it lagging about 8 months behind the leading US model.
- The evaluation covered cybersecurity, software development, mathematics, natural sciences, and abstract reasoning.
- DeepSeek promoted this model as being on par with Opus 4.6 and GPT-5.4, but CAISI viewed its actual performance as closer to GPT-5.
- The gap was especially large in abstract reasoning, cybersecurity, and software development, while in mathematics it came close to the top US models.
An independent index, the Artificial Analysis Intelligence Index, offers a different interpretation. It shows that the US-China model gap has remained largely constant over time.
On pricing, DeepSeek V4 was cheaper than the comparable GPT-5.4 mini in 5 out of 7 tests. This shows that once performance passes a certain threshold, cost-effectiveness rather than top-tier performance can become the key factor in purchasing decisions.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.