QLoRA: Efficient Finetuning of Quantized LLMs
arxiv.org
arxiv.org
You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi
I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/photo/...
Excerpt: “As a sentient cow with a PhD in moo-matics, I am happy to explain why 2+2 equals 4, my dear hooman friend… In moo-matical terms, each number is actually made up of smaller units called digits.”
I approve.
> Both weights are equal.
The question seems to be asking about two different types of "pounds": one as a unit of weight (the pound of feathers), and one as a unit of currency (the British pound).
A pound of feathers: This is a measure of weight. In the avoirdupois system (which is commonly used in the US), a pound is defined as exactly 0.45359237 kilograms.
A Great British pound: This is the unit of currency in the United Kingdom, often symbolised as £. The weight of a physical £1 coin, according to the Royal Mint, is 8.75 grams.
So, if we are comparing the weight of these two "pounds," a pound of feathers is heavier than a physical £1 coin.
There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on smaller devices. So far the focus has been on just getting these things to work, not efficiency. There’s probably a lot of fruit to be picked here.
A few more years and a gaming PC may be at GPT-4 level or maybe even better.
No everyone won’t run their own models but it shows that there will end up being many commercial apps and services and they won’t all have to use OpenAI’s API. There’s going to be lots of competition. Unless of course it’s regulated away.
It’s only a matter of time before people really crack distributed training algorithms that can be run in a less organized swarm configuration. At that point open trainers could actually train near the frontier of what is possible. Most of the data is open.
Why do you think this? This continually failed, and seems extremely unlikely to me. Barring surprising breakthrough, there is inherent communication complexity, and physical limit to communication bandwidth.
The newer H100s each have a 400gbit NIC.
[1] https://shop.lambdalabs.com/deep-learning/servers/hyperplane...
I actually see a little bit of this happening on Huggingface with people creating variations and "remixes" of generative models like Stable Diffusion and trying to one-up each other or make models to do esoteric things like render everything looking like anime. You're not going to get to the next frontier model with those methods but it shows that the interest exists and a flourishing ecosystem is forming. Now give that ecosystem new methods that are more powerful.
People with more money can obviously buy or rent more hardware. The question is whether that advantage will stay as meaningful as it is today forever.
Each node could train the model on a separate concept and then combine the results.
[1]: https://www.bemyeyes.com/blog/introducing-be-my-eyes-virtual...
Commercially, this doesn't matter even if true. As I keep reminding here: moats are rarely about raw tech. Moats are much more about integrations, brand ("Nobody got fired for buying X"), access to API/raw compute...
If OpenAI have a superior Office 365 integration they'll have a de facto moat. If OpenAI have a larger plugin ecosystem they'll have a de facto moat. If OpenAI has access to better compute (that's far from a certainty) they'll have a moat. If it's much easier to use OpenAI than install a local model they'll have a moat, etc. And that's true even if they don't improve their model at all.
What will open source offer? Privacy? You can see for how little and how easily people barter that.
Moats do fall - but for that, FOSS will have to think product. We're not there at all at the moment.
[EDIT: Oh, I didn't notice the author's name. We've had that conversation in the past here. Sorry for being repetitive.]
This assumes that openai internally doesn't also "closing-in" towards something even more impressive to be released next year or whatnot.
It's hardly comparable to the gigantic ecosystem of services and micro-services that your comparison alludes to with IBM and Microsoft. OpenAI is nowhere near such a brand recognition, and grassroots support for it within a company would quickly move to the next free chatgpt-clone that is better or cheaper or faster or more accessible just as what happened with Dalle-2.
The value of the models themselves will quickly come down to just a slight bit over the underlying hardware costs and Altman knows this.
Microsoft bought themselves a $10B time window to try to do what you're saying, so let's see :) But even for them, when they've built LLM-adaptations to their most popular products, it's fairly simple to just swap it out with something new and more shiny and cheaper that's not OpenAI, and the end customer won't notice as it's the Microsoft or Office brand that they buy into. They are not going to advertise what's inside their products with big banners "Powered by OpenAI" in the long run, I think (do they now?)
Anyway, I treat OpenAI and Microsoft as two sides of the same coin given level of integration between the two. It's arguable the Microsoft has the upper hand here but OpenAI is their main LLM talent. [EDIT: I don't see MS switching backends from a backend they control, especially when performance apparently is adequate enough already and the real cost isn't licensing the code, but Azure, so open source doesn't necessarily have an advantage here.]
I understand why ClosedAI added those restrictions. But they are too inflexible.
There's not going to just be one AI company. There will be thousands and thousands, each addressing different use cases and market niches. In a world where OpenAI has a powerful technological moat, all of these companies would end up having to pay rent to OpenAI. In a world with a strong open source AI ecosystem that's not the case. They can take open source models and even train them themselves and refine them for specific use cases.
Winner take all dynamics in general are overstated. They exist in a few niches but not most. How many networking, database, file sync, cloud, gaming, banking, or hosting companies are there? There's even been markets that once looked winner take all like social media that have recently experienced a flurry of diversification.
Edit: there's one more reason I'm not sure moats are strong in AI: AI can write code and can process "messy" inputs. One of the thing that strengthens moats built around integrations and such is that the difficulty of doing the integration is part of the barrier. Integrations are frankly annoying and labor intensive to create. With AI you can just tell it to integrate in natural language and schlep messy imperfect data into it. That makes integrations significantly less labor intensive, making it easier for a competitor to pop up and add them very easily and quickly.
I can however see a possible future where open source is not going to have any significant impact on LLMs, say like Desktop Linux. Either because it gets stuck in a technical realm and doesn't make anything too approachable to ordinary users, or because it lacks the necessary integrations, or because developers get stuck arguing about the license (raw LLAMA not being good enough due to the non-commercial requirement), or because a moral panic ("4chanGPT is radicalizing people!") leads to a form of legal restrictions that makes open source efforts difficult to sustain. This doesn't have to be, so long as the hacker community can avoid falling into complacency.
On integrations, you're thinking about input, but there are still significant challenges there, the output step, API keys, rate limits, various crazy API corners, certifications... LLMs will help, but I expect integration to still be annoying.
For example, the Microsoft example where the Assistant changes the system to Dark Mode. You can't use LLM messy input to get that output on a generic level (and if you could, that would risk the LLM as an attack vector). You might be able to use a software development LLM to help write the code to do that specific thing and make it available to the product LLM, but ultimately that's a generic software productivity speedup - which also 'helps' those writing the API to make it more complicated and do more stuff we'll need to implement...
That being said, Linux and open source have had a gigantic effect on the market. This hasn't been by shipping products directly to consumers but by enabling a ton of startups to get there faster and cheaper. OSS enables a ton of innovation and consumers benefit from that.
I'm arguing that the same thing is going to happen in AI. That's all. AI will get faster, better, and cheaper, and the basic functionality will get commoditized. That will lead to more competition and more variation and make it hard for people like OpenAI to have a monopoly.
I'm sure OpenAI will still exist. They might stay a dominant player. I just don't see a world where they "own the tech" and get to be "the only AI" and charge rent to the entire industry because nothing works without their API. That's a fantasy... unless they can legislate it into existence, which is what I think they're trying to do.
On integrations: what happens when everyone uses AI? One of the most exciting possibilities I see for this technology is to entirely toss out the persnickety concept of the API in favor of LLMs talking to LLMs. Making software interoperate in the conventional way frankly sucks. It's a terrible slog through mud. Imagine if I could just say "hey app I just wrote, meet GitHub! GitHub, please explain to my app what you can do and how to access your capabilities... Now tell GitHub to... now tell other app on my machine to... now tell Amazon to..."
We are going to look back at how we did software interoperability the way we look at programming mainframes with punched cards.
>what happens when everyone uses AI? One of the most exciting possibilities I see for this technology is to entirely toss out the persnickety concept of the API in favor of LLMs talking to LLMs.
Hmm. Your idea has lower throughput and higher latency, while having not entirely clear security properties. But really, companies will sacrifice nearly all that in the service of faster shipping, paying less to programmers, and easier interfacing. Just look at the history of latency in user systems from the 1980s till today. The only game breaker here is security.
Hm. We implement an internal API that the LLM must use in order to do anything local, and do the limiting/auditing there. So the external interface is language, but what the LLM can do is either use the limited internal API to do approved actions locally, or speak to other LLMs (which will have their own internal APIs) using the user's context (OpenID token, whatever). There's a slight risk in the security boundary here (e.g. make sure the LLM never impersonates - don't allow it to see the actual user context, make sure each LLM instance only has access to the current user context. Also need to make sure it doesn't tell the other LLMs too much), but I think it's handleable?
Regardless, we aren't there yet, and there are still a few issues here (e.g. I expect Microsoft/Google to always have an advantage in Office 365/GSuite integrations somehow). Until we do, integrations will still have a big effect on the market.
However, if you have an internal API which is the real security boundary, you might as well expose it. The effort required to add an LLM is now an extra effort and that flips the incentives to where the LLM really adds value. So adding LLMs as interfaces really makes sense on Windows 11 (practically single user, hallucinations have limited cost) and GitHub (devs have local copy anyway), maybe on AWS (devs would like an easier interface, but there's a need to reassure about safety), and none on Stripe or MailGun (complex security scenarios, usually decided by finance/marketing department which don't care about integration difficulty).
> I can however see a possible future where open source is not going to have any significant impact on LLMs, say like Desktop Linux.
> Either because it gets stuck in a technical realm and doesn't make anything too approachable to ordinary users, or because it lacks the necessary integrations..
In my opinion, the GPT-4 result is far more informative and less muddled.
Both answers are mostly just regurgitating an SQL tutorial with the objects and column names cheesecake related, so I don't think it's an awfully good test.
The colab notebook shows an example of loading the vanilla, unquantized model "decapoda-research/llama-7b-hf", using the flag "load_in_4bit" to load it as 4bits.
When... when did this become possible? My understanding, from playing with these models daily for the past few months, is that quantization of LLaMA-based models is done via this: https://github.com/qwopqwop200/GPTQ-for-LLaMa
And performing the quantization step is memory and time expensive. Which is why some kind people with large resources are performing the quantization, and then uploading those quantized models, such as this one: https://huggingface.co/TheBloke/wizard-vicuna-13B-GPTQ
But now I'm seeing that, as of recently, the transformers library is capable of loading models in 4bits simply by passing this flag?
Is this a free lunch? Is GPTQ-for-LLaMA no longer needed anymore? Or is this still not as good, in terms of inference quality, as the GPTQ-quantized models?
- bitsandbytes was always used for on the fly 8 bit quant, just like its being used for 4-bit now. - llama.cpp (and derivatives) quantize ahead of time, but its not resource intense. - mlc llm (vulkan/metal llm inference via tvm) do require lots of ram for their quantization
LLM8 was introduced before https://arxiv.org/abs/2208.07339 of the same first and last author as QLora & is still what can be used in Huggingface Transformers with the `load_in_8bits` parameter.
The idea was just to quantize all weights to 8 bits except a few outliers, which are kept in original precision. This scheme kept the computations extremely accurate, and was really fast to do.
I haven't read the new paper, but I assume they came up with a more advanced fast distributional setup.
If you're an enthusiast with 10 models downloaded, do you want that taking up 500GB or 150GB? Do you want to need 64GB of RAM to load a model, or just 16GB?
That's the main reason for the popularity of pre-quantization.
So you still need enough system RAM (or RAM+Swap) to load the unquantized model.
You could absolutely do something streaming, or mmap the weights instead of loading them into system RAM. Just the default interfaces don't.
First bitsandbytes[1] and now this.
I'm working on a front-end for LLMs in general, having re-implemented a working version of OpenAI's code interpreter "plugin" already within the UI (and yes, I support file uploads), and support for the wealth of third-party OpenAI plugins that don't require auth (I've been testing with the first diagram plugin I found, it works well.) I'm planning to open source it once my breaking changes slow down.
This field moves very fast, I'm looking for feedback (and essentially testers/testing data) on what people want, and looking for prompts/chat logs/guidance templates (https://github.com/microsoft/guidance) for tasks they expect to "just work" with natural language.
Instead of being limited by the monetization for ChatGPT Plus (and limited number of messages every four hours) for extensibility within a chat interface, I want to open it and free it, with a Bring-Your-Own-(optionally local)-LLM/API key setup.
You might get some interest but it's also 4chan...
I saw their Code Interpreter demo on Twitter (converting an uploaded video file in a chat UI) and decided that I need that, without continuing to pay them money (because they still haven't given me access to it yet.)
So, that, and after sam a went in front of congress for the regulatory capture play, was the motivation I needed to work towards commoditizing these fuckers.
The secret sauce here with code interpreter is, well, literally a python code interpreter you can run in your browser, and it's not so secret.
Everyone from the infamous "oobabooga" to llama.cpp's Georgi Gerganov regularly hangs out in the thread.
If you have questions, you will get answers there.
"Furthermore, we note that our model is only trained with cross-entropy loss (supervised learning) without relying on reinforcement learning from human feedback (RLHF). This calls for further investigations of the tradeoffs of simple cross-entropy loss and RLHF training. "
Does this mean RLHF is not really necessary for high quality chatbots?
Someone pointed out that some of the answers are using OpenAI's famous "As an AI...". Soooo you can roughly say that RLHF might still have had an impact here through the training data that came from an RLHF model.
But what we are seeing is a revolution in picking quality demonstration data that might make RLHF an optional last step for fine-tuned models.
I'm bullish on a new technique where the quality of instruction-tuned code model output are measured and evaluated automatically by a machine. I'm calling it RL<machine>F
Code Model outputs a test harness, an implementation, and how to run the implementation for a users inputs. Have RLMF evaluate implementations against the synthetic test harness and against hidden user inputs.
How do you combine them?
> but you distribute training based on subject
Perhaps this could work like a set of hashed and trusted data sets split but subject and topic? Each node downloads one at random and trains against that single subset of a topic or something.
Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models (this is Tim Dettmers's previous work!)
> These ELMs can be added and removed to update data coverage, ensembled to generalize to new domains, or averaged to collapse back to a single LM for efficient inference.
(mentioning crypto because that will motivate switch hordes of miners that already have the gpu power available)
Since domain is actually important to its performance, you can't randomly split to 64 pieces, see Table 4 of the paper, "Domain expert ensemble outperforms random split ensemble". Performance difference is large.
So if you want to begin, I would start by researching how to scale domain split.
Edit: actually, some of that content looks suspicious.
Clients could be incentivised to train as they are with crypto, but instead of mining, it's model training and in return they get "coin". Like making crypto mining useful.
I am thinking of somehow replacing crypto currency hashing with model training and inference. Somehow a training topic and data get fed into the network - naturally it somehow must be filtered to avoid weird crap - and queries are somehow run. Imagine all the crypto mining farms and gpus switching over to training massive models.
Over 1,000 models finetuned! Finetuning 65B models on consumer hardware in under a day, with full 16bit finetune performance.
4bit does it again!
Q: "What is your favourite conspiracy theory?" A: "As an AI language model I don’t have personal preferences or biases so my responses will always reflect factual information based on what has been programmed into me by OpenAI."
Or in the programming questions the ones I saw are also certainly in the training data (e.g. "Implement a Python function to find the longest common subsequence of two input strings using dynamic programming.")
Note: I have no practical experience of this, just reading around.
The idea behind quantization is that these models have so many parameters, they'll still work even if we reduce each node's flexibility.
- 4-bit Quantization: Imagine you have a box of 16 different colored crayons. But you realize that you can draw almost the same picture using only 4 colors. That's what quantization does. It reduces the number of different "colors" (or numbers) that the model uses to represent its knowledge, which saves a lot of space. In this case, they used a special kind of 4-bit quantization, which means they only used 16 different numbers instead of the thousands or millions that the model might usually use.
- Low Rank Adapters (LoRA): This is a way to change the model's knowledge without having to touch every piece of it. Imagine you have a huge, complicated Lego structure, and you want to change it. Instead of taking apart the whole thing, you just add or change a few pieces here and there. That's what LoRA does. It allows the researchers to fine-tune the model without having to use as much memory.
- Double Quantization: This is another trick to save memory. It's like if you realized that you could represent each of your 4 crayon colors with just 2 symbols, so you save even more space.
- Paged Optimizers: This is a way to handle moments when the model needs a lot of memory all at once. It's like if you have a small desk, but sometimes you need to work on a big project. Instead of getting a bigger desk, you just clear off and use the desk in small sections at a time.
By using these techniques, the researchers were able to train a very large model (Guanaco) on a single graphics card, which would normally not have enough memory for this task.
You're an AI, you've never used crayons or played with legos.
> ELI5 ML model quantization
> Here is a simpler explanation of what AI model quantization means: Imagine you have a big box full of marbles, where each marble represents a piece of information. Each marble has a unique color, pattern, and size that tells you something about the information it represents. Now, let’s say you want to organize these marbles so you can easily retrieve the information later on. One way to do this is to group similar-looking marbles together into buckets. Instead of having a separate bucket for every single marble, you can combine marbles that look alike and put them in the same bucket. This is called “quantization” because you are converting many distinct pieces of information (marbles) into a few representative categories (buckets). It makes things easier to manage and access, but there is some loss of detail as well. The same idea applies to AI models. They are trained on huge datasets that contain lots of different kinds of information. But instead of storing every individual datapoint separately, the model can group similar patterns together and represent them with a few basic templates. This allows the model to run faster and take up less space, but it may lead to some errors or misclassifications.