LoRA Learns Less and Forgets Less
arxiv.org
arxiv.org
LoRa has been a popular wireless protocol for like 10 years.
(1) well, except Google which -surprise- returned about mid page an ad of a local quite expensive chandeliers brand called "LORA".
99% of ML engineers wouldn't know what it is.
A much smaller percent among those who write in ML (the functional programming language) probably, though.
It's virtually impossible to be sent for a loop if you're engaging critically. The two domains solve very different problems and have no overlapping concepts. The way we talk about each topic is very distinct.
And yet (as others have noted) adding another word or two to give a bit of context is usually enough that web searches work.
Search engines are pretty clever nowadays, except when they've been deliberately dumbed-down (cough... Google...).
Isn’t an equally valid argument that MLPs tend to constitute a greater number of weights in transformer networks than attention heads, and the performance difference can be traced to a greater number of weights having freedom to change? I’d be curious to know if randomly choosing a subset of matrices to train, regardless of where they are in the network, would provide analogous performance to LoRA on a specific module with comparable learnable weights.
Galore might be more equivalent to full pretraining with the gradients being low rank.
For example, tuning a layer of 128 in x 256 out is 32k params. Learning a full-rank lora for that layer would be two matrices of 128x128 and 128x256 = 48k params.
Other papers show finetuning a select few layers can also work well.
In a continual learning paper from last year, I found LoRA was extremely effective for faster fine-tuning and not forgetting the original dataset:
Not to be confused with LoRa ("long range") [1], an Internet of Things radio technology.