Dust: Pretraining Transformers Without Backpropagation (qlabs.sh)

64 pointsby E-Reverance2 hours ago3 comments

polyomino an hour ago

Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory

api an hour ago

It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?

vatsachak an hour ago

Not necessarily, backprop is highly parallelizable since it is just a bunch of matrix mults.

Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".

The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.

Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params