Q Labs just dropped Dust — a wild new way to pretrain transformer LLMs WITHOUT backpropagation. Yeah, you read that right. No gradients.

How it works:
→ Injects random noise into layer outputs at every token
→ Rewards changes that cut the loss
→ Tested on models from 2M to 243M parameters

Early results? It trails backprop at first, but the gap closes hard as you scale to tens of millions of tokens. Bigger models = more efficient use of the method.

Code is open source. This could flip how we think about training efficiency and compute costs in AI infra.

If this scales to billion-param models, we're looking at a potential paradigm shift for $AI infrastructure tokens and compute marketplaces.