LoRA: Low-Rank Adaptation of Large Language Models
github.com
github.com
I believe there will soon be a cottage industry of providing application-specific fine-tuned models like this, that can run in e.g. AWS very inexpensively. The barrier today seems to be that the base model (here, Meta's LLaMA) is encumbered and can't be used commercially. Someone will soon, I'm confident, release e.g. an MIT-licensed equivalent and we'll all be off to the races.
[0] https://www.databricks.com/blog/2023/03/24/hello-dolly-democ... [1] https://twitter.com/alighodsi/status/1639251347777388544
https://huggingface.co/EleutherAI/gpt-neox-20b
Perhaps it’s in the paper…
https://huggingface.co/serpdotai
I haven't done any formal tests on this yet, but with llama-13b, the overall structure of its responses definitely becomes much more ChatGPT-like. It would be very interesting to see how the 65B model performs.
>it needs a plugin for ControlNET
The big difference is that ControlNet actually required a pretty complex interface to be used effectively, meanwhile the use of LoCon/LyCORIS should be completely transparent and works like a LoRA
Reproduced is a strong statement, without any rigorous justification other than a few cherry-picked examples. Alpaca-LoRA is simply LLaMA with LoRA-tuning on the Alpaca data. There are no metrics, no measurements, no evaluations to show that the Alpaca-LoRA performs similarly to Alpaca, when it is well-known in the field that parameter-efficient fine-tuning always pays a cost in terms of performance relative to full fine-tuning (which is what Alpaca does).
(This has been a huge nit for me because of the recent flood of Alpaca-replications, or even claims that Alpaca comparable to ChatGPT, rushing to market themselves, but with nothing to justify their claims.)
It also bothers me that a lot of LoRA claims read like "You won't believe how little it costs to train these models!", when of course 99%+ of the complexity and cost is in the LLaMA (or whatever) model that underpins it. Folks are talking about it in a loose way that implies some kind of miraculous overall training cost breakthrough.
The LoRA paper clearly states the performance of the method "LoRA performs on-par or better than fine-tuning in model quality on RoBERTa, DeBERTa, GPT-2, and GPT-3, despite having fewer trainable parameters, a higher training throughput, and, unlike adapters, no additional inference latency. ": https://arxiv.org/abs/2106.09685
As simple way to think about it is this: if LoRA really gives full fine-tuning performance, why would anyone ever fully fine-tune a model?
You're asking it as if it were a rhetorical question, but I think it carries more weight than many people seem to believe.
That said, I also dislike it when it is carelessly claimed that parameter-efficient tuning is as good as full fine-tuning, without qualifications or nuance.
It is not apparent to me that fine tuning should be better, especially since the LoRA method seems like it could be robust against catastrophic forgetting.
But to be clear, LoRa is not related to ANN training, is it? Why would they be using both?
On their github they reference a related project called "HuggingFace" so you know the sky's the limit with the names in this field, could have been called anything else really.
Quick jargon literacy boost: "HuggingFace" is a platform tailored to hosting and sharing ML repositories -- like Github for AI. The parent company, "Hugging Face", is also in and of itself a major contributor to several AI research projects & tooling.
Ironically, they still managed to hit a namespace collision... albeit self-inflicted.
I think you’ve illustrated the problem the lawyers were protecting them from.
Star Wars: Dark Forces
Star Wars Jedi Knight: Dark Forces 2
Star Wars Jedi Knight 2: Jedi Outcast
Star Wars Jedi Knight: Jedi Academy
See also: Huggingface PEFT: https://github.com/huggingface/peft
Original weights frozen = Rather than modify the original model, the training results are saved to a small file of only a few MB.
In practice this means you can fine tune a 30B parameter model on a consumer GPU in a couple of hours. Without LoRA you would need to run multiple expensive data center GPUs for days or weeks.
(1) https://en.wikipedia.org/wiki/Low-rank_approximation
Edited: By the way, it seems to me that there is an error in the wikipedia page because if the Low-rank approximation takes a larger rank then the bound of the error should decrease, and in this page the error increases.
It seems that the initial matrix of weights has a low rank approximation A and this implies that the difference E = W - A is small, also it seems that PCA fails when E is sparse because PCA is designed to be optimum when the error is gaussian.
Since the weights are derived from gradient descent, yeah we don't really know what the distributions would be.
A random projection empirically works quite well for very high dimensions, and is of course very cheap computationally.
Think of a point cloud of a piece of paper floating in the wind. It would be a 3xn list of points, but "really" it's a 2d piece of paper.
Just like I can rewrite the number 27 as 333 or 8+19 or (2^3)+(2^4)+3.. Given a single matrix one can find myriad ways to rewrite it as a sequence of matrices that have the same (or similar) numeric value, but with interesting or desirable properties. :D
My favorite example (which is used in signal processing) is to take your ugly matrix and rewrite it as a set of smaller matrices where most of the elements are zero, or a power of 2.
It turns out, computers can multiply by zeros and powers of two very fast
Why is fine-tuning done with separate alterations, rather than by mutating the original weights?
It's actually larger. If you just have two equally large matrices of the same dimension, one original, and one of "altercations"... then you can just add them together.
> Why is fine-tuning done with separate alterations, rather than by mutating the original weights?
Then you'd have to compute the gradients for the whole network, which is very expensive when the model has 7b, 65b, 165b parameters. The intent is to make that cheaper by only computing gradients for a low rank representation of the change in the weight matrix from training.
You have to do that with LoRA regardless, to compute the gradients for the lowest-level LoRA weights.
The goal of most parameter-efficient methods is to store one gold copy of the original model, and learn minor modifications/additions to the model. The easiest way to think about this is in some kind of deployment setting, where you have 1 capable model and you learn different sets of LoRA weights for different tasks and applications.
The original intent of parameter-efficient methods is to reduce the amount of storage space needed for models (do you really want to keep a whole additional copy of LLaMA for each different task?). A secondary benefit is that because you are fine-tuning a smaller number of parameters, the optimizer states (can take up to 2x the size of your model) are also heavily shrunk, which makes it more economical (memory-wise) to (parameter-efficient) fine-tune your model.
From the LoRa paper:
>When the pre-trained model is GPT-3 175B, the number of train- able parameters |Θ| can be as small as 0.01% of |Φ0|.
Consumer GPU, yes, but in practice LoRA doesn't actually reduce training time. What it mainly reduces is memory requirements. In fact LoRA training can often require more training steps than full fine-tuning and therefore be slower (you can imagine why this is the case: the optimization is trying to modify the mode's behavior a smaller number of parameters, and so has a harder job)
Here's an example of 20 seconds per epoch on a single consumer GPU: https://github.com/johnsmith0031/alpaca_lora_4bit/issues/7#i...
Can you qualify this? Is it still useful or not?
Is what still useful? A LoRA is about as good and useful as a full fine tune. If you have unlimited storage space to store them or unlimited compute to make them then I would still prefer full fine tunes. But the difference is marginal and generally not worth the storage space or increased compute costs for individuals.
The insight is that we don't need to modify a lot of parameters to get a generally competent model to do well on specific tasks. When you have a linear layer with a weight matrix of dimension d_in x d_out, the change you undergo during full finetuning is also a matrix of d_in x d_out, which can be huge. We represent the latter using two matrices of shape d_in x r and r x d_out. You save a lot of parameters when r is small. So when you use it, the input goes through two streams 1) the orignal frozen weight turning a vector of size d_in to d_out and 2) the low-rank weights turning a vector of size d_in to r and r to d_out. The two streams are then summed together. (There's a figure in the paper.)
This way of doing thing is nice for a few reasons. It's easy to parallelize. You can change r to control how many parameters to train. You can also merge the low-rank weights with the original one to avoid latency.
Note that we don't select a subset of the original parameters. We train extra ones.
I am assuming you can have n LoRA fine-tunings, say each specializing in one aspect of a coherent task, with n summers, running in parallel, and then combine them at the end? Or more generally, does LoRA enable a sort of modularizing around a core (un-merged) model?
And curious if you ever tried merging 2 or more fine-tunings and then testing the resultant single model (merge all) against the original tests to check retention?
https://arxiv.org/pdf/2202.13914.pdf
The gain isn't that significant. We don't understand what these low-rank updates represent, and they might not correspond to "skills" that humans have.
My background contains signal processing, "pre-deep learning ML", systems engineering, and firmware, and that sentence jumped out at me as crystal clear in my mind, despite not knowing what HuggingFace is or PyTorch.
Correct me if I'm wrong: These huge models involve lots of weights used in large matrices. The contribution of this work is to plug in some matrix factorization and learn a lower dimensional representation, instead of a large second matrix.
Fantastic!
Also makes me wonder what other performance improvements await through proper application of established and well known Mathematics. :D
[0] https://towardsdatascience.com/adding-custom-layers-on-top-o...
Add a task specific layer and only training that layer doesn't work well. In practice, people combine many of these things, e.g., LoRA + task-specific final layer.
Not exactly the same, to be sure. But fulfills a similar need: more efficient "fine tuning" of a large model.
"The other direction, as exemplified by prefix tuning (Li & Liang, 2021), faces a different challenge. We observe that prefix tuning is difficult to optimize and that its performance changes non-monotonically in trainable parameters, confirming similar observations in the original paper. More fundamentally, reserving a part of the sequence length for adaptation necessarily reduces the sequence length available to process a downstream task, which we suspect makes tuning the prompt less performant compared to other methods."
https://ar5iv.labs.arxiv.org/html/2106.09685
This is key imo: "More fundamentally, reserving a part of the sequence length for adaptation necessarily reduces the sequence length available to process a downstream task".
The benefit of prompt and prefix tuning (note: these are two separate methods) is that you can serve different soft-prompts and soft-prefixes efficiently with a single shared set of model weights.
> incurs a non-trivial computation cost
The hit seems to be in energy/cpu not time since the W0 computation is in parallel with the BAx. (My assumption based on the latency claims in paper.) So an issue in edge deployments (battery life, etc.).
> you are stuck with that one model on that device
Upfront I have 0 clue on the actual numbers, but from a purely software architecture pov [in unmerged setup], having that W0 forward process once with n distinct BAx paths (for distinct fine tunings!) would address that, no?
[p.s. say an application that takes as input A/V+Txt, runs that through an Ensemble LoRA (ELoRA™ /g) which each participant contributing its own BAx finetuing processing, sharing the single pre-trained W0.]
The latency claims are based on the merged version, where the modifications are merged into the model weights. Hence there is no latency cost, since the final model has the same shape as the original.
> having that W0 forward process once with n distinct BAx paths (for distinct fine tunings!) would address that, no?
The tl;dr is that that works, but is more expensive. Not ridiculously more expensive, but certainly more expensive that processing a few additional tokens with prefix/prompt tuning.
If one is careful with floating point issues, it's straightforward to unmerge the weights.
W_0 = W_1 - BA
Yes, prompt-based methods don't involve swapping weights.
(Also, I'm assuming you're the first author of LoRA.)
Yup, I am!
Prompt tuning does so by injecting addition prefix tokens in the input to the model. LoRA does so by injecting low-rank matrices that are additive modifications to a set of linear layers in the model.
They both do something slightly different, but are very much in the same class of methods.
Edit: Microsoft is even a member of the LoRa alliance: https://lora-alliance.org/lora-alliance-press-release/micros...
https://en.m.wikipedia.org/wiki/Lora
It doesn't really matter as long as it's not in the same field. No one will be confused between the two.
Except search engines...
I bet some of them have even been to Minnesota and they still didn't pick a unique name.
Though both of them have to answer to why they picked the name of a Google Font that preceded both and is currently available https://web.archive.org/web/20170210001724/https://fonts.goo...
Is it because Microsoft is competing with Google in the AI space?
Not the same situation, but I remember when “Electron” was called “Atom Shell” because it was built for the (now defunct) text editor by the same name. For the longest time, I had an unsubstantiated thought that it was a new Unix shell that was based around a text editor somehow (yes, dumb). In hindsight, they just had named this cleverly to reference the various layers or shells of electrons orbiting atomic nuclei, thus the eventual name of Electron.
On the other hand, a wireless technology standard is very different than a known mathematical technique that likely predates the wireless meaning anyway.
This process involves low rank approximations -> Lora is a namey sounding term that uses characters from low and rank -> call it LoRA in the paper. That’s all there was to it. Probably didn’t even know the other lora existed.
That said some of these ML project names are especially horrendous (kind of ironic for the current emphasis on generative AI). Transformers? A good chunk of the time I get results about the toys and cartoons from my childhood. Don’t get me wrong, I still think Optimus Prime is cool and the name “transformers” make sense given the function but it’s somehow simultaneously generic AND the name of a decades long multi-billion dollar media franchise…
LoRA is another example, name makes sense but the collision with LoRa is problematic. I, for one, am interested in and have/would apply both. Queue google searches for “Lora radio…” vs “Lora ml…”.
Project naming is hard and I’m just glad to see the activity and releases. BUT project naming is essentially a base usability condition and should be considered as such: just like creating a README, getting started, providing code examples, etc.
It reminds me of trademarks: if you’re looking for trademark protection it won’t be issued if it is overly generic or likely to “cause confusion in the marketplace” with an existing trademark (basically same or similar name in a somewhat similar/adjacent field) - you can even reuse names but only if it’s obvious to people from basic context that they refer to different things. I’m not a trademark attorney but I think LoRa vs LoRA would get refused because it’s “computer stuff”, while a shampoo named Lora would be fine (as an example). If you’re curious there are official categories/areas from the USPTO that break these down.
Both of these examples wouldn’t have a chance at trademark protection. Note I’m not saying they should have trademark protection, just that it’s an example of a reasonable standard that should be considered/compared to for good open source project naming.
The LoRA paper’s ‘problem statement’ makes a compelling case for practical benefits of the approach. Specifically, no added latency, no serial processing bottlenecks, shared baseline model, compact time/space requirements. How does dreambooth stack up in this regard?
“Aghajanyan et al. (2020) shows that the pre-trained language models have a low “instrisic dimension” and can still learn efficiently despite a random projection to a smaller subspace.”
Would be great to have an informed practitioner comment (sota) on why we opt for random projection. Is the actual ‘intrinsic’ vector space uncomputable? Too slow to find?
https://en.wikipedia.org/wiki/Random_projection
> The core idea behind random projection is given in the Johnson-Lindenstrauss lemma, which states that if points in a vector space are of sufficiently high dimension, then they may be projected into a suitable lower-dimensional space in a way which approximately preserves the distances between the points.
OMG yes. It's also an unsolved problem, once can approximate it. I've written several non-parametric blind arbitrary dimension DR algorithms and been obsessed with the space most of my life. If you think O(n^2) feels slow, try O(n^3)...:) For more, read about mean shift clustering, or the new hot stuff: topological data analysis/bar codes/mapper algorithm.
There's already a technology called LoRa!
Fuck I hate this crap. Be better than this.
Unlikely that they even consider checking whether they are stomping across existing names.
Or it's on purpose as existing terms already have good amount of search traffic for those terms, and Microsoft know Google/Bing will rank Microsoft's own pages higher than what's already out there.
/s