AI Briefing
KO

Magic Announces Pretraining Research Achieving Performance Similar to DeepSeek V4 Pro Base with Approximately 50x Fewer FLOPs

·2026.09.08 09:00

Key point

Magic announced that a new pretraining recipe reached performance similar to DeepSeek V4 Pro Base with approximately 50x fewer FLOPs, and that training scaled up by approximately 10x subsequently outperformed public base models in perplexity evaluations.

Details

The Magic team announced research results for a new pretraining recipe evaluated as more than 10x compute-efficient compared to major open-weight base models. This recipe reached performance similar to DeepSeek V4 Pro Base with approximately 50x fewer FLOPs (approximately $500,000 based on GB200). Subsequent training with approximately 10x the compute (approximately $4 million) showed significantly superior performance in perplexity evaluations compared to all public open base models. Model release was presented as a future plan.

Compute Efficiency and Performance Comparison

Magic measured bits per byte (bpb) loss on heldout data and fit scaling laws to estimate the compute required to reach the same loss. Lower bpb indicates better results, and the following are bpb measurements for V5 e24 and DeepSeek V4 Pro, along with the estimated efficiency multipliers of the new recipe in that domain.

  • Reasoning (Math): 127x efficiency compared to DeepSeek V4 Pro (0.587 bpb vs 0.678 bpb)
  • Private Code Repos: 48x efficiency compared to DeepSeek V4 Pro (0.194 bpb vs 0.202 bpb)
  • Heldout Research Papers: 45x efficiency compared to DeepSeek V4 Pro (0.383 bpb vs 0.404 bpb)

These multipliers are not simple ratios of bpb values or actual training FLOPs ratios between V5 e24 and comparison models, but comparisons of compute required to reach identical performance using fitted scaling laws.

Verification Method and Data Contamination Prevention

A rigorous verification process was applied to objectively evaluate the model's generalization ability. To prevent data contamination, they used their own private codebases, CoT-based math problems generated by Kimi K3, and recent low-citation research papers, removing documents that exceeded 96-character normalized text matching or Jaccard similarity thresholds. Additionally, documents were rewritten by third-party frontier LLMs to ensure that sequence memorization was not rewarded.

Future Plans and Roadmap

The Magic team, one of the smallest teams in the world training large-parameter models, plans to focus on long-horizon RL scaling, long-context learning after deployment, and developing alignment techniques for narrowly induced latent knowledge. They also announced an update to the AGI Readiness Policy, including deployment gates and safety requirements during RL training.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.