Dust Enables Transformer Pretraining via Zeroth-Order Node Perturbation
October 5, 2026
Dust approximates backpropagation by perturbing activations independently at every token during a single forward pass. The method scales more efficiently with model size, with a 243M-parameter model showing higher population efficiency than significantly smaller architectures. This approach offers a compute-heavy alternative to gradient descent that may potentially surpass backpropagation performance in specific regimes.
HOW THIS AFFECTS YOU
●
builderThis provides a theoretical path to training models without standard backpropagation infrastructure.
●
researcherYou can explore training paradigms that bypass the need for differentiable architectures.