AI Briefing
KOSign in

DeepSeek V4.1 Flash Leads Agentic Coding Benchmarks, Narrowing US-China AI Gap to 3%

·2026.10.07 00:44

Key point

DeepSeek's new model scores 77.3 on agentic coding tasks, surpassing Anthropic's 66.1, but carries legal risks under China's National Intelligence Law.

Details

The performance gap between top US and Chinese AI models has narrowed to a record-low 3%, driven by DeepSeek's release of V4.1 Flash. According to a Bloomberg Intelligence analysis citing the October 4 LiveBench snapshot, DeepSeek now outperforms Anthropic's leading model on agentic coding, a critical benchmark for enterprise software automation. DeepSeek scored 77.3 on this sub-benchmark, compared to 66.1 for Anthropic's Claude Fable 5.1 Max Effort.

Architecture and Pricing

DeepSeek V4.1 Flash utilizes a Mixture of Experts (MoE) architecture with 748 billion total parameters, including a new persistent memory component called Engram. Only a sparse subset of these parameters activates during inference, allowing for significantly lower computational costs. The model is priced at $0.30 per million input tokens and $1.20 per million output tokens at peak, with off-peak rates at half that cost. It supports a context window of up to 1 million tokens and adds native multimodal capabilities.

Benchmark Discrepancies

While LiveBench data shows a narrow gap, other methodologies yield different results. NIST's Center for AI Standards and Innovation (CAISI) evaluated the predecessor model, DeepSeek V4 Pro, in May 2026 and estimated the US-China gap at approximately eight months of development time. Analysts note that different benchmarks measure different aspects of performance, and the 3% figure reflects a composite leaderboard rather than parity on all specific enterprise tasks.

Legal and Security Risks

A major constraint for enterprise adoption is China's National Intelligence Law, which requires Chinese companies to cooperate with government intelligence requests. Data sent to DeepSeek's API, including proprietary code and prompts, is subject to potential disclosure to Chinese authorities. This legal obligation cannot be overridden by privacy policies. Additionally, security researchers previously identified exposed databases at DeepSeek, raising operational security concerns.

Enterprise Implications

For teams requiring strict data governance, self-hosting the open-weight model on private infrastructure eliminates API transmission risks but requires significant hardware investment. The US government has already banned DeepSeek on government devices, and bipartisan legislation aims to formalize these restrictions. While DeepSeek leads in raw agentic coding performance, its ecosystem maturity and integration tooling lag behind US competitors like Anthropic and OpenAI.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.