Takeaways from hundreds of LLM finetuning experiments with LoRA
lightning.ai
lightning.ai
The only thing I wish you did different was explored alpha > 2 * r.
In this blog post, the author found that alpha of 4 * r (where r=64) outperformed all smaller alphas in terms of loss when finetuning Llama-7b on databricks-dolly-15k.
https://medium.com/@drishtisharma96505/comparative-analysis-...
Additionally you identify (alpha = 2*r) r=16 as inferior to r=256, however aside from arithmetic, r=16 actually outperforms all others. And the base model outperforms any finetuning for both arithmetic metrics.
This experimental process on small NNs us exactly how science is done. You develop and test theories of what will work using small networks because you can do them quickly.
Then you test on large, production size networks based on the lessons from the small ones.
It's possible that does not work. LoRA (for Low Rank) benefits from the "small changes" introduced during finetuning of a model. The update of the weights has a low rank. If you take a smaller model, it might induce that the rank is not so low, resulting in degradation in metrics by LoRA compression. I would be interested to see if LoRA still has a benefit in this configuration.
Pedantic, but it is actually for Low Rank Adaptation.
Microsoft already created this huge confusion, so I'd recommend you help not make it worse by using the relevant capitalization.
LoRA is LOw Rank Adaptation, the LLM tuning technique
Even if not in this specific comment thread, maybe having correct capitalization here means someone else knows how to write the one they mean to talk about in another conversation later, and then it will have been helpful.
If everyone went around calling both of them LORA all the time absolutely that's going to sometimes have people talking past each other about different technologies. Even with them capitalized correctly it's still going to, but not as much.
Too bad Microsoft didn't call theirs LoRaA and it would've been more obvious.
They (and I) may need to do a double-take if they process it wrong, and you reduce that risk by being as clear as possible, despite Microsoft's incredibly poor naming choice.
By analogy, if I refer to MICRO soft, and their operating system, I'm sure there's no confusion to you. But it is harder to read, and may (for longer sentences) require re-reading the whole sentence.
(and, of course, the other replies you got, and other people being confused every single time there's any announcement or article about LoRA)
LoRAs can be nearly the same size as the original model, with nearly the same representation capacity/trainability, or they can be a tiny tiny fraction of the original model size, with correspondingly fewer learnable parameters.
As such, they are suitable for nearly all tasks. We should be asking if they are better than regular fine-tuning, or soft-prompts (aka textual inversion), or slapping new trainable layers on the end (aka hypernetworks). The stable diffusion community seems to think that they are.
> My hypothesis is that the Alpaca dataset does not contain any related arithmetic tasks, and the model actively unlearns basic arithmetic when it focuses more on other tasks.
I'm surprised this wasn't verified, it's a major benchmark stat. My eyes keep getting drawn to it, because it seems to have the most variance. Does anyone know?
Also throwing it out, I would love to see a Neptune/Wnb of the hyperparameter tuning :)
Aka "catastrophic forgetting" (CF).
The most popular benchmarks and datasets are mostly haphazardly cobbled together with hardly any oversight or verification, sometimes even synthetically generated from gpt 3.5 et al. without checking the output at all. Frankly it's amazing that any of it even works when people blindly train and test with what's essentially self contradicting garbage.
>> As a result of an accident, Abdul lost sight in his right eye. To judge the distance of vehicles when he is driving, Abdul is able to rely on cues of >> - A. I only >> - B. II only >> - C. III only >> - D. I and II only
> You didn’t read that wrong. The question never explains what I, II or III are. This appears to have been improperly copied from crackap.com. Somehow Platypus 2 still gets the right answer with high confidence. Is this a sign it has merely memorized the answers? I checked the second best ranked model upstage/LLama-2–70b-instruct-v2 and it also somehow got the answer right (the third best Open LLM also gets this question right so I don’t know what is happening).
https://derenrich.medium.com/errors-in-the-mmlu-the-deep-lea...
The article explores optimizing LoRA hyperparameter settings for finetuning large language models. The goal is to maximize performance while minimizing memory usage and training time.
The base model used is Llama-2 7B. Experiments compare default LoRA, QLoRA (4-bit quantized), AdamW, SGD optimizers, and different choices for rank r and alpha hyperparameters.
Key findings:
QLoRA provides substantial memory savings (6GB less than default LoRA) at the cost of slower training. Performance impact is minor.
AdamW vs SGD makes little difference in memory or performance.
Increasing training iterations from 50k to 100k hurts performance, likely because the Alpaca dataset lacks diversity.
Tuning rank r and alpha is most impactful. Good rule of thumb is to set alpha=2*r. Best model uses r=256, alpha=512. Improves over base model on most tasks, except arithmetic.
The optimized LoRA model was submitted to the NeurIPS efficiency challenge and showed improvements on several benchmarks compared to the base Llama-2 model.
Takeaways are practical tips for tuning LoRA hyperparameters and trading off memory, compute, and model performance.
I was VERY confused for a minute.
But on top of that, this is also one of the many fun examples of where a static, linear document that's so simple it could have been written in markdown requires megabytes of download across dozens of files.
Nobody cares about user experience right now -- it's all about developer and content management experience, so runtime bulk and responsiveness are ignored in favor of... whatever these kinds of developers think they're gaining on the production end. I often have a hard time guessing what that is without assuming the responsible devs are not-so-great.