I wonder if you could take a traditional backprop trained LLM and apply this approach to finetuning it (presumably needs less memory and compute?). It could be another entry in the spectrum between LORA and full fine tuning.
Turns out this is more compute intensive than regular backprop. The win (if any) would be in training on future highly distributed architectures where neighbouring parts of the model have accesss to very high bandwidth (to each other) but communicating to further away parts is more expensive