AI Briefing
KOSign in

Dust: First Zeroth-Order Method Competitive With Backpropagation in Transformer Pretraining

·2026.10.06 06:15

Key point

Dust introduces a zeroth-order optimization method for transformer pretraining that perturbs activations per token, achieving competitive loss with backpropagation and significantly outperforming weight-space Evolution Strategies like EGGROLL.

Details

Core Innovation: Activation-Space Perturbation

Dust is a zeroth-order optimization algorithm that replaces backpropagation with activation-space perturbation. Unlike traditional Evolution Strategies (ES) that perturb weights, Dust adds independent Gaussian noise to the output of linear layers for each token. This creates a virtual population where each token acts as an independent population member, allowing a single forward pass to evaluate a large population in parallel. This approach decouples population size from the number of forward passes, enabling Dust to assess populations at least three orders of magnitude larger than weight-space ES methods like EGGROLL within the same computational budget.

Performance vs. Backpropagation and EGGROLL

In experiments using GPT-style transformers on the FineWeb dataset, Dust demonstrated that zeroth-order methods can be competitive with backpropagation:

  • Efficiency: From 1 million tokens up, Dust is estimated to be 10^3 to 10^4 times more efficient than a transformer implementation of EGGROLL, a state-of-the-art weight-space ES method.
  • Loss Metrics: At 100k and 1M token budgets, Dust achieved lower test loss than backpropagation with small populations. At larger budgets (10M and 20M tokens), the gap with backpropagation narrowed significantly as population size increased, with fits suggesting the gap continues to close.
  • Comparison to ES: Dust significantly outperforms weight-space ES methods in terms of efficiency and loss reduction for a given compute budget.

Scaling and Gradient Alignment

Contrary to the conventional belief that zeroth-order methods fail on large networks due to variance, Dust showed that larger models are more population-efficient:

  • Model Size: A 243M parameter model performed similarly to or better than a 2M parameter model across most population sizes, challenging the idea that zeroth-order methods are limited to small networks.
  • Gradient Alignment: As population size increases, the cosine similarity between Dust’s estimated gradients and true backpropagation gradients improves, following a fitted law: cos(K) = c_max / sqrt(1 + c/K). This alignment holds across different layer types and token counts up to 1B tokens.

Limitations and Future Directions

While Dust approximates backpropagation closely and is competitive in loss, it is not yet a practical replacement for backpropagation due to compute efficiency constraints. The authors note that Dust requires orders of magnitude improvements in compute efficiency to compete with backpropagation at current scale. However, by removing the requirement for end-to-end differentiability, Dust opens new avenues for architecture search, including recurrent and looped computations that are difficult to train with standard backpropagation.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.