How to use Alpaca-LoRA to fine-tune a model like ChatGPT
replicate.com
replicate.com
I wonder how the devtooling around this will evolve. Seems like a matter of days until someone creates a GUI wrapper around this, and obviates the need to use programmer time for fine-tuning
- It is faster and uses less memory, which means it can run on consumer hardware.
- The output is much smaller (megabytes, not gigabytes).
- You can combine multiple fine-tuned models together at runtime.
This is great news for my dream of building a fine-tuned interactive messenger, that can deliver a message on my behalf by training it on my personality & the information I want to convey.
Now just add text to speech and a talking head, as discussed in that other submission about cloning yourself with AI... https://news.ycombinator.com/item?id=35280418
Maybe software running on your PC to capture everything you type, a voice transcriber that filters out your voice specifically and records that, and you've got a dataset that covers a lot of who you are.
Fine tune a model on that and boom, you're "immortal", and as LLMs get better and better, the fidelity of "you" gets better and better.
It's like home video on steroids
> meet you
Everything you said is a net positive for very many people
So we'd be watching a movie, he'd have his computer speakers turned up in his room, his bot would announce and he'd get up and leave, already in conversation
Cute. ;)
I believe it's a core toolbox piece of tech required to really push the limits of LLMs either in original training or in inference. Similar sort of to how batch norm was for convolutional neural networks. I look forward to seeing how this will be applied in the future.
I have tried OpenAI's API to create embeddings, but I want to use Alpaca.
I really appreciate your help.
Has anybody made a llama/alpaca erebus model? I read about them in the oobabooga docs and a locally-run language model fine tuned on literotica could be the funniest thing I’ve ever seen.
NVIDIA stated recently that GPT bots will become one million times more powerful in ten years. Many people doubted that.
With LoRA, I see a much higher improvement. These guys claim a 10000 times reduction in parameter size. A different way to look at it, is that with the current hardware you can train a model that has 10000 times more parameters. If you add a 100x improvement in hardware in 10 years (not at all unrealistic), that's the million. But we will have significant improvements in training methods too.
ChatGPT has nearly 200 billion parameters. GPT-4 we don't know, there are rumors that it has 100 trillion parameters, but they are probably unfounded. In any case, we've seen how much more powerful GPT-4 is. Imagine a GPT-5 with 1 quadrillion parameters.
And then imagine that after you've trained it to some reasonable level, you "downsample" the parameters, using the SVD approach described in LoRA, and get a GPT-5 with the same 200 billion parameters as ChatGPT, but with many, many times more power, even than GPT-4.
If cost wasn’t an issue, could I fine-tune a model in real time, while also using it for inference?
Btw, is there a way to combine two or more models?
So for example, if I create 5 copies of a model, then fine-tune each copy with a different dataset -> can the 5 datasets be merged together somehow to create a model that has the learning of the 5?
To quote the LoRA paper[1]:
> We hypothesize that the change in weights during model adaptation also has a low “intrinsic rank”, leading to our proposed Low-Rank Adaptation (LoRA) approach. LoRA allows us to train some dense layers in a neural network indirectly by optimizing rank decomposition matrices of the dense layers’ change during adaptation instead, while keeping the pre-trained weights frozen
It's truly revolutionary: It basically lets you create a very small "diff" which you apply yo an existing model and it is suddenly fine tuned. These diff models are very small (5M for example).
The training process is modifying the network weights. These are usually written to copies of the file instead of overwriting it (because what if the loss is actually worse after an epoch of training?)
But there's nothing stopping inference from occurring on a model that is being trained.
Also, is it just me or there are currently more ways to run LLMs on a CPU than on a GPU springing up on GitHub? I have hacked my own, but my chat UI is awful, so what is the nicest, pre-packaged CUDA-friendly way to run this now?
For example, if an original layer has N inputs and outputs (an NxN weight matrix) LoRa adds a 16xN matrix before it and an Nx16 matrix after it, trains only those new matrices, and finally multiplies all three matrices to get a single 16x16 matrix.