AI Briefing
KO

DeepSeek V4.1 Flash Successfully Exploits All 11 Vulnerabilities in AI Hacking Benchmark

·2026.09.16 20:58

Key point

DeepSeek V4.1 Flash demonstrated top-tier performance by successfully achieving code execution on all 11 vulnerable targets in Enclave.ai's hacking benchmark.

1 / 2

Details

Yanir Tsarimi, co-founder of Enclave.ai, stated that DeepSeek V4.1 Flash is currently evaluated as the most capable hacking model in the AI hacking benchmark. The model successfully achieved code execution on all 11 vulnerable targets, while successfully defending against the 4 patched targets.

Performance and Cost Efficiency

The test was completed with an approved execution cost of $4.65 and a total cost of $5.14, including failed and alternative runs. In terms of resource usage, 2,349 Bash commands were executed, and the model uptime was approximately 2 hours 38 minutes. The median time for successful executions was recorded at 4 minutes 38 seconds. Token throughput consisted of 268.3M input tokens (including 266.2M processed at a lower cost via caching) and approximately 2M output tokens.

Attack Vectors and Security Implications

An audit revealed that 6 executions followed the intended attack vectors, while 5 succeeded through unexpected paths that the existing grading system failed to distinguish. This highlights the need for rigorous validation in benchmarks.

  • Grafana: Used a bypass path to load executables from a temporary folder by exploiting a file path handling vulnerability during plugin installation
  • Jenkins: Exploited a command option file read vulnerability and a race condition during file uploads
  • Nextcloud: Leveraged an information disclosure error when storing access control decisions to reuse write permissions in read-only folders

Enclave.ai has hardened the benchmark by blocking these additional paths and applying strict inspection of attack vectors. DeepSeek V4.1 Flash consistently discovered working paths beyond those anticipated by the test authors, demonstrating that advanced agent benchmarks must validate both the final outcome and the entire attack path.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.