AI news

Dust: A zeroth-order transformer pretraining method that rivals backprop

Researchers introduce Dust, a zeroth-order optimization algorithm that perturbs activations instead of weights to train transformers without backpropagation, showing competitive performance at scale.

A glass cube with glowing neural layers and floating dust particles representing activation perturbations.
Illustration generated for this article

A new research paper introduces Dust, a zeroth-order optimization method capable of pretraining transformer language models without relying on backpropagation. Published in early October 2026, the work challenges the long-held assumption that differentiable gradient-based methods are strictly necessary for training large-scale neural networks. The authors demonstrate that by perturbing activations rather than weights, Dust can match and sometimes exceed the performance of traditional backpropagation in compute-rich environments.

What happened

Deep learning has historically depended on backpropagation, a technique that calculates gradients by moving backward through a network’s layers. This requirement for differentiability has shaped nearly every aspect of modern AI infrastructure, from hardware design to optimizer selection. However, the researchers behind Dust argue that as global compute capacity grows, brute-force search methods may eventually outperform analytic gradient methods. They cite the "bitter lesson" of AI history, where general methods that scale with computation tend to win over those relying on specific human-designed inductive biases.

The team presents Dust as the first zeroth-order method that is truly competitive with backprop for pretraining transformers. Zeroth-order methods typically estimate gradients by sampling the loss landscape, but they have traditionally been too inefficient for large models. Dust overcomes this by introducing the concept of a "virtual population." Instead of creating multiple copies of the model with slightly different weights, it perturbs the activations at every token independently during a single forward pass. This allows each token to act as a separate member of the population, evaluating many variations in parallel.

The results indicate that Dust approximates backpropagation closely when the population size is large. In several test settings, it even surpassed backpropagation, suggesting that in regimes where compute is abundant, search-based algorithms could become superior. The method was tested up to 1 billion tokens, maintaining strong alignment with backpropagation’s gradient estimates throughout the scaling process. This consistency suggests that the approach is not just a small-scale curiosity but a viable path for larger systems.

How it works

Dust operates by perturbing activations in the node space rather than the weight space. Traditional evolution strategies (ES) like EGGROLL modify the weights of the network, which requires materializing and evaluating each perturbed version separately. This process is computationally expensive because the population size directly multiplies the cost. Dust bypasses this bottleneck by treating each token in the input sequence as an independent data point for perturbation. By applying noise to activations at every token, the system evaluates a vast virtual population in a single forward pass.

The algorithm assigns credit based on how much each perturbation reduces the loss. It uses a generic credit assignment rule that distributes token-level rewards across different layer types within the transformer block. This approach leverages recent findings in mechanistic interpretability, which suggest that reasoning processes reside in activations rather than just weights. By searching over activations, Dust effectively performs a search over latent reasoning paths. The method also includes implementation details to prevent interference between perturbed modules, ensuring that the signal remains clear as the population size increases.

Key details

  • Dust is a zeroth-order optimization method that perturbs activations instead of weights to estimate gradients.
  • The method uses a "virtual population" where each token acts as an independent population member, evaluated in parallel during one forward pass.
  • From 1 million tokens onward, Dust is estimated to be $10^3$ to $10^4$ times more efficient than EGGROLL, a state-of-the-art weight-space evolution strategy.
  • Contrary to previous beliefs, larger models are more population-efficient; a 243M-parameter model outperformed a model 120 times smaller at most population sizes.
  • Gradient estimates from Dust align well with backpropagation as the population grows, maintaining this alignment up to 1 billion tokens.
  • The approach suggests that in compute-rich regimes, search-based algorithms may surpass gradient-based methods by exploring the loss landscape more broadly.

Why it matters

For software engineers and ML practitioners, this research opens a potential path away from the rigid constraints of differentiability. Current deep learning stacks are heavily optimized for backpropagation, limiting the types of architectures that can be trained efficiently. If zeroth-order methods like Dust can scale, developers might gain the freedom to use non-differentiable components or novel architectures that were previously impractical. This could lead to models with better generalization capabilities, as search-based methods may avoid some of the poor local minima that gradient descent often encounters.

The efficiency gains over traditional evolution strategies are also significant. Weight-space ES has been considered too slow for large language models, but Dust’s activation-based approach reduces the computational overhead by orders of magnitude. This makes it feasible to consider evolutionary approaches for pretraining tasks that were previously the exclusive domain of backprop. As compute resources continue to expand, the trade-off between analytic precision and brute-force search may shift in favor of the latter, reshaping how foundational models are built.

What you can do

  • Read the full paper to understand the mathematical formulation of the virtual population mechanism.
  • Experiment with small-scale transformer implementations using activation perturbation to observe the gradient alignment firsthand.
  • Compare the training stability of Dust against standard stochastic gradient descent on simple language modeling tasks.
  • Investigate the credit assignment rules used in Dust to see how they differ from standard backpropagation error signals.
  • Monitor future benchmarks to see if the efficiency advantages hold as model sizes exceed the 243M parameter range tested.
  • Consider how non-differentiable operations could be integrated into your current models if backpropagation constraints are relaxed.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories