AI Briefing
KO

Magic AI Releases 'Magic V5 e24', Achieving Performance Similar to DeepSeek V4 Pro with 48x Fewer Computations

·2026.09.08 09:00

Key point

Magic AI has released the Magic V5 e24 model, achieving performance similar to DeepSeek V4 Pro with approximately 48x fewer FLOPs.

Details

Magic AI has released the Magic V5 e24 model, announcing that it achieved more than 10x computational efficiency compared to existing open-weight models. Notably, it achieved performance similar to DeepSeek V4 Pro with approximately 48x fewer FLOPs, which is half the computational volume of GPT-3.

Computational Efficiency and Performance Metrics

Magic V5 e24 demonstrated overwhelming efficiency compared to competing models across various domains, including Private Code, Research Papers, and Math Reasoning. Lower bits per byte (bpb) metrics indicate better performance, and V5 e24 recorded lower bpb than major competing models such as DeepSeek V4 Pro, Kimi K2, and Nemotron 3 Ultra.

  • Private Code: V5 e24 (0.194) vs DeepSeek V4 Pro (0.202, 48x efficiency)
  • Research Papers: V5 e24 (0.383) vs DeepSeek V4 Pro (0.404, 45x efficiency)
  • Math Reasoning: V5 e24 (0.587) vs DeepSeek V4 Pro (0.678, 45x efficiency)

Verification Methods and Data Contamination Control

To ensure evaluation reliability, logprobs were verified through vLLM, GB200, and the Fireworks engine. To prevent overlap between training and evaluation data, documents with 96-character matches or high Jaccard similarity were removed. Additionally, a process was undertaken to rewrite or summarize documents using third-party LLMs to prevent sequence memorization. The AIME benchmark was excluded due to concerns about external model contamination, and a private heldout competition math eval (400 problems) was used instead.

Reinforcement Learning (RL) Performance and Future Plans

Without SFT or distillation, performing math RL at the base model stage with a 16k CoT budget resulted in V5 e24 achieving Pass@1 72% on heldout competition math. While this falls short of GPT-6 Astra (100%) and Claude 5.1 (98%), it is at a level similar to Kimi K3 (75%). Notably, the e24 model broke through AIME26 pass@1 90% when using only 0.2% of the pretraining compute budget.

Magic AI plans to continue improving pretraining efficiency while maintaining the world's smallest team size for training trillion-parameter models, alongside scaling long-horizon RL and post-deployment learning.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.