Dust: A zeroth-order transformer pretraining method that rivals backprop
Researchers introduce Dust, a zeroth-order optimization algorithm that perturbs activations instead of weights to train transformers without backpropagation, showing competitive performance at scale.






