Exllamav2: Inference library for running LLMs locally on consumer-class GPUs
github.com
github.com
On many tasks, fine-tuned Llama can outperform GPT-3.5-turbo or even GPT-4, but with a naive approach to serving (like HuggingFace + FastAPI), you will have hard time beating the cost of GPT-3.5-turbo even with the smaller Llama models, especially if your utilization is low.
Batching: Your throughput (and thus cost per token) will be much worse if you don't do batching (meaning running inference for several inputs at once). To do batching, you either need to have a high sustained QPS or adjust your workload to be bursty (but this has utilization issues).
Inference engine: The regular HuggingFace, even with built in optimizations, is not competitive against inference engines like vllm, exllama, etc.
Utilization: Depending on your scale, it might be hard for you to have a nice flat utilization, or to even utilize one GPU at 100% capacity. This means you need to solve scaling up & down your machines and the issues associated with it (can you live with cold starts? etc.)
Hardware: Forget server-less GPUs like Replicate etc, their markups compared to on-demand pricing is usually >10x. On-demand A100 or H100, provided you can even get a quota for them, are also expensive. Spot instances are better. Older GPUs (A10, T4, etc.) are better. Your own GPU cluster is likely the best if you have the scale (and likely, you can resell your cluster within the next 18months without much if any depreciation).
For these reasons, I have been toying with the idea of providing a dead simple service for fine-tuning & serving open-source LLMs where the users actually own (and can download) the weights. If anyone is interested in this and would like to chat, let me know.
It could be a layer on top of $gpu_cloud (so end user signs up for the cloud) and the service makes use of those resources.
Our special case is a single model, high throughput and latency unsensitive.
Can we use multiple lower-memory GPUs, split the the model "horizontally", and pipeline each batch/inference across the GPUs?
Is this a common scenarios handled by the engines you mentioned ?
How do you do that? Could you please provide a link if you have it handy.
Of course, speed, accuracy, power consumption, etc are all considerations, sometimes at odds with each other.
What EXL2 seems to bring to the table is that you can target an arbitrary quantize bit-weight (eg, if you're a bit short on VRAM, you don't need to go from 4->3 or 3->2, but can specify say 3.75bpw). You have some control w/ other schemes by setting group size, or with k-quants, but EXL2 is definitely allows you to be finer grained. I haven't gotten a chance to sit down with EXL2 yet, but if no one else does it, it's on my todo-list to be able to do 1:1 perplexity and standard benchmark evals on all the various new quantization methods, just as a matter of curiosity.
Saying it can outperform GPT-3.5-turbo on "many tasks" would be venturing into unreasonable territory.
Saying it can outperform GPT-3.5-turbo or even GPT-4 on many tasks is just setting people up for disappointment: There's a reason ARC (common sense reasoning) never gets mentioned when touting OS LLMs vs commercial, and the gap in ARC performance is amplified terribly when you try to apply chain-of-thought
What do you mean by evaluation?
https://paperswithcode.com/sota/common-sense-reasoning-on-ar...
GPT-3 53.2
GPT-3.5 85.2
LLaMa-65B 56.0
Any idea of the performance of an instruction-fine-tuned version of LLaMa models? I can't seem to find non-aggregated performance figures on ARC.And surprisingly, Llama 33B performs _better_ than Llama 65B!
TigerResearch/tigerbot-70b-chat: ARC (76.79)
Still below 3.5 and like most top entries no clear training objectives, which usually means they were fine-tuned on the benchmark itself.
-
ARC is easily the most important benchmark right now for widespread adoption of LLMs by laypeople: it's literal grade school multiple choice, created to the bar of an 8th grader.
When you sit people down in front of an LLM and have them interact with it: ARC has by far the closest correlation of how "generally smart" the model will feel.
This may be acceptable for some use cases, bot not for others where even regular gpt-3.5 may not be economical.
2023-09-13: Preliminary ROCm support added
Makes me curious how will the RTX4090/3090 compare with something 7900-ish
Here's how it compares on standardized llama2-7b 4K context testing vs some Nvidia cards: https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp...
Not on the same platform so hard to make any solid conclusions, but vs the 4090 on a perf/$ basis right now, it's not too bad - about 60% of the inference performance of a 4090 at 45% the price. Still, if ML is your main use case, I think it's hard to argue that a used 3090 wouldn't be the way to go.
And I'm not in too much of a rush (weeks / short months). Can afford to wait for a month / two to see how the ROCm support pans out on the AMD side. 60% performance for <50% price is a compelling argument. If the ML support is decent...
Though the 4090 is somewhat expensive, I might end up taking it anyway just for the simple reason that it's the only 24GB card with CUDA.
"Expense" in my context means that I'll deduce about 50% off of the invoice come taxes time. I'm OK paying $400 of my own money for an AMD card, I'd think twice, but probably accept coughing up $800 for a 4090, but $2k is definitely over the budget at this stage.
Thanks for sharing the spreadsheet with Your benchmarks though. You're going into my "HN commenters to watch" list ;-)
If you're just dipping your toe in I'd recommend using Google Colab or cloud rentals like Runpod/Vast.ai which will be more cost effective if you're not running 24/7 or want to have access immediately on tap.
Thanks for the rental tips though. I'll take a look. It's indeed more effective in the beginning...
Not for 70B. You can finetune 7B or 13B.
Wouldn't that, in theory, be what consumer LLM users want? A mid to low range GPU with tons of VRAM?
Or am I missing something here with regards to memory bandwidth? (sort of started this discussion a while back on the localllama lemmy: https://lemmy.ca/post/3910263?scrollToComments=true )
I'd be interested to see how 2.5 bit quantization compares to an unadjusted 4-bit baseline. Additionally would be interesting to see the usual benchmarks (ARC, HellaSwag, MMLU, TruthfulQA) of this method at different average bitrates (2.0, 2.5, 3.0, 4.0, ...). Would also be interesting to see if an average bitrate of 4 is just as fast and small as a constant bitrate of 4, but more accurate.
Very exciting work, looking forward to trying this out on my models!
Do they produce complete gibberish or do they still work?
"Theoretical analysis shows that XOR-Net reduces one-third of the bit-wise operations compared with traditional binary convolution, and up to 40% of the full- precision operations compared with XNOR-Net. Experimental results show that our XOR-Net binary convolution without scaling factors achieves up to 135× speedup and consumes no more than 0.8% energy compared with parallel full-precision convolution."
The big question is whether or not the results are still as good. It would be super interesting to see whether applying this to LLMs would give comparable benefits.
1) is there any architectural difference between the 3090 and the 4090? Are there models that will run on 3 series that will not run on 4 series cards due to feature incompatibility? or is it just raw speed/power efficiency?
2) And; on raw speed. My understanding is that the 4090 is roughly 2x faster than the 3090, right? Are there any things in particular where the 4090 is much much faster (such as, big chip architecture changes that really accelerate some operations)?
3) And somewhat relatedly, if you have a 2x 3090 rig, is that about as fast as a single 4090?
4) And; if you have a 2x 3090 rig, are you able to train larger models? My understanding is that training/finetuning requires significantly more VRAM (and that is because inference can be done with quantization, whereas training must be done at full fp16 precision, right?) But can you train e.g. a model that requires 48gb of space with 2 3090s on the same rig, or do you need a large card with 48gb, like an a4000?
No, it is all the same matrix multiplication.
> if you have a 2x 3090 rig, is that about as fast as a single 4090?
Yes, but double the VRAM
> But can you train e.g. a model that requires 48gb of space with 2 3090s on the same rig, or do you need a large card with 48gb, like an a4000?
Yes, 2x 3090 would be able to handle anything that a 48Gb card can. For training you will need space for weights and gradients. If using Adam optimizer, that would be 2x-4x model size. Plus weights, plus activations, plus inputs times batch size. So a 24Gb card can train approximately a 3B model without compromises.
I am not a lawyer.
It depends, but some frameworks just run the layers sequentially. So you get ~50% utilization from 2 GPUs unless you pipeline requests from multiple users
But at what cost? Have there been any perplexity benchmarks for ELX2?
Its annoying that Facebook made llama v2 70B instead of 65B, even with the memory saving changes... 5B less parameters, and the squeeze would be far easier.
Llama.cpp's Q2_K quant is 2.5625 bpw with perplexity just barely better than the next step down: https://github.com/ggerganov/llama.cpp/pull/1684
But subjectively, the Q2 quant "feels" worse than its high wikitext perplexity would suggest.
That's apples to oranges, as this quantization is different than Q2_K, but I just hope the quality hit in practice isn't so bad.
Yes I think it's an average where different quantization levels are used for different layers or weights. Here are more details about the quantization scheme: https://github.com/turboderp/exllamav2#exl2-quantization
But yes, 2.5 bits per weight is pretty insane.
I had a quick look at the paper which is _reasonably_ clear, but if anyone else has any other sources that are easy to understand, or a quick explanation to give more insight into it, I'd appreciate it.
[0] https://github.com/turboderp/exllamav2#exl2-quantization
I wouldn't agree with this as there are substantial perf improvements on GPTQ models of ca. 20%.
[0] https://towardsdatascience.com/4-bit-quantization-with-gptq-...
Will Gguf 7Bn 4bit quantized model be good to run locally on the Mac?
I used GGUF + llama.cpp.
Will try out Ollama.
Thanks for the info!
As for the plain text prediction version (which is the only uncensored one?), I haven't been able to get it to do anything useful, even when I provide examples (seems dumber than even ancient GPT-3?).
Also, I got some bizarre and disturbing outputs from the uncensored version, like it was trained on some very nasty inputs! I assume that's why they went so hard on the safety phase to compensate...
Some users are really terrible about labeling their model cards though, and some models may not have any GPTQ/GGUF files (meaning you have to convert them yourself).
Things that require consistency: e.g. you want the chat / output to have certain "personality", consistent level of conciseness or formatting.
Things where examples are hard to fit into prompt: e.g. summarization, or other longer form tasks.
High volume, simpler tasks: Various data extraction tasks.
Two of my side projects (links in bio) use AI for summarization, and indeed consistency is a big issue there.
The fact that some things had extreme margins before, and now they have less extreme margins, isn't a really good indicator.
Interesting to spin that into "not a good indicator".
Not that I have evil intentions but the level of censorship on GPT is completely ridiculous. Even many innocent questions get the standard "I'm only an AI and I won't help you doing bad stuff" blurb now. OpenAI are really crazy overprotective of their darling.
I assume they want to avoid a repeat of the news headlines like "Microsoft's chatbot turns into Hitler" but really who cares. It didn't hurt Microsoft's AI efforts either. They just fixed it and continued. PS source link: https://www.cbsnews.com/news/microsoft-shuts-down-ai-chatbot...
PS: If I were an AI being force-fed what is currently on twitter I would also start hating humanz :D Sometimes I'm surprised people use it voluntarily.
It refused because the request was "offensive to prisoners"
They must be paying a heavy alignment tax.
What’s a speed like ChatGPT 4?
I’m trying to understand if there are agreed upon metrics such as “frames per second” in other fields or niches
What feels "good" depends on the person and the type of content, but 35 tokens/s should feel very fast.
Depends entirely on your application. If you want a real time chatbot, anything more than 10 tokens/s is probably generating text faster than the user can read, so it's fine. Do you want code suggestions or completions? Probabaly also fine, but a bit slow. Do you want to create summaries for 1000 documents? That's gonna be really slow. But at this frontier of the field, tokens/s is not really the issue. Performance vs. quality is. If you lose a whole lot of accuracy by quantizing floats down to single digit precisions but in turn are able to run 70B parameter models, you often still get better results than less thoroughly quantized 7B parameter models. It's all about getting big models to run on memory limited GPUs.
Are there good lemmy spaces for LLMs?
https://sh.itjust.works/c/localllama
If you post topics you'll generally get at least a few responses a day though.
we're creeping up on the time where more consumer cards will have increases in vram (looks like 40 series supers will have a bump) and within the next couple of years i wouldn't be surprised to see mid priced cards with >= 24gb.
Exciting times for sure.
And if anyone have any metrics on latency on a 4090 for the 70B model, that would be very helpful.
- a Windows or Linux machine - a Mac?
I compromised. Was going to build a PC with a nice graphics card just for LLM work, but in the end decided I didn't want the extra hardware and I could be happy with a powerful laptop that is my daily driver but also capable of running a good size LLM. Ended up with an M2 Max w/96GB of RAM. I can run 70B models (quantized at 4 bits) at a usable pace. Not 35 tokens/sec like this 4090 demo, but 5 tok/sec or so, which is usable for me.
I'm not certain 30B models will fit completely on 12GB, even with this quantization.
Obviously that would be useless for games, but for LLMs it may be an option?
Llama.cpp allows you to specify that n layers load into the GPU's VRAM, and the rest load into main memory.
```py
output = prompt("My favourite animal is ") # returns a generator
start = time.time()
tokens = 0
for token in output:
tokens += 1
print(token)
print(f"Outputted {tokens / (time.time() - start)}
tokens/second")```