RC RANDOM CHAOS

Dust: A Breakthrough in Training Transformers Without Backpropagation

· via Hacker News

Original source

Dust: Pretraining Transformers Without Backpropagation

Hacker News →

Researchers have developed Dust, a zeroth-order method for pretraining transformer language models that rivals backpropagation in efficiency. Dust perturbs activations at each token, treating every token as a virtual population member, allowing for parallel evaluation. The method approximates backprop closely with substantial compute and even outperforms it in multiple settings, suggesting that brute-force search methods could surpass backprop in high-compute regimes.

Dust is significantly more efficient than weight-space evolution strategies (ES), being orders of magnitude faster than state-of-the-art methods like EGGROLL. Contrary to common beliefs, larger models are more population-efficient with Dust, indicating that overparameterization might provide a better search space. The alignment of Dust’s gradient estimates with backprop improves with larger populations and remains consistent up to 1 billion tokens, making it scalable for future applications.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.