AI Briefing
KO

DeepSeek Achieves GPT-5.2-Level Performance

·2026.05.05 15:51

Key point

DeepSeek V4 Pro delivered GPT-5.2-level performance on agentic benchmarks while costing 17 times less.

Details

In testing by FoodTruck Bench, an agentic performance measurement tool, DeepSeek V4 Pro entered frontier-model territory, posting results on par with GPT-5.2.

Key findings from the analysis are as follows:

  • Performance and Ranking: It ranked 4th overall. While its results were similar to Grok 4.3, it held an edge in execution consistency, such as reducing food waste and increasing the number of meals served.
  • Overwhelming Cost Efficiency: It delivers equivalent performance at roughly 17 times lower cost than GPT-5.2 on agentic workloads. In terms of cost efficiency (net worth relative to API spend), it ranked 2nd overall, behind Gemma 4 31B.
  • Narrowing Technology Gap: The gap between US and Chinese models, once around a year, has now narrowed to roughly 10 weeks.
  • Chinese Models on the Rise: Xiaomi MiMo v2.5 Pro also placed 6th on the benchmark, showing that lower-priced Chinese models are proving highly competitive in the agentic market.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.