High-Speed Large Language Model Serving on PCs with Consumer-Grade GPUs
github.com
github.com
After some thoughts, in ReLU it does make sense, because half of the function is constant, so you can say that you're "cold" if that neuron's ReLU-ed output is often 0 . So I checked whether ReLU was common in LLMs, original llama doesn't use ReLU. But after (re-)reading the github, it actually only works on ReLU models. Turns out that there is a group of people "fine-tuning" (I would rather call that re-training, since you start by breaking the model?) models to use ReLU to allow for that sparsity: https://huggingface.co/SparseLLM
So this is sadly not applicable to any model you can find on the internet, but that sounds like a great progress anyway. Possibly this might shift the compromises back to bigger models but with "less ideal" activations. Also I'm curious what would be the legal impacts on it (since USA and EU refers to a model's FLOPs/number of parameters... How do you compute it with sparsity? Do you average?)
I think that a possible avenue for future research in that area is keeping original activation (like llama keeping SwiGLU), but using quantification to define "hot" and "cold" neurons to be saturation areas. (For example, saying that this activation function, below -1. at 8 bit, is equivalent to -infinity, and thus this is a cold neuron)
https://huggingface.co/SparseLLM/ReluFalcon-40B
"We utilize PowerInfer for inference"
How/when did these types of regulations come about? This feels like an insane thing to have to keep in mind while developing.
They're trying to get in early on AI so as not to make the same mistake again. Which might result in them making the opposite mistake.
Making the world a worse place? If you look carefully you’ll realize most of the harms and negative effects of technology are due to it being primarily funded by advertising and trying to maximize ad revenue.
I don’t think mobile gaming companies have a potential to destroy free press, or negatively affect mental health of wide population of teenagers, or invade privacy of billions of people. They simply don’t have the scale for any of that.
'I watched an ad, and then [my entire life was destroyed]' is quite hard to imagine, unless it's an ad for an MLM, crypto, entrepreneurship scam, or gambling.
On the other hand, I absolutely know people who started out in soft gambling who then proceeded to throw their life (and sometimes families) away trying to catch the next high with higher and higher stakes gambling until they lost everything, and then some.
We also don't really know the impact gambling is going to have in the near future. Loot boxes, online gambling, internet celebrity gambling, etc. really only became popular around ~2010 or later, and the kids who have been growing up with low-risk gambling as a daily accessible thing on their iPads have not come into adulthood yet.
It is still unethical to even play "free"-to-play games. You are entertained at the expense of a small group of addicts that are often spending more money than what they can afford, and, at least in many games, just being logged in helps create a nicer environment that lures in those people. If you are not there to be a whale you are there to be lure for them. It might not be harmful to you to play, but you are being harmful to the addicts.
The biggest benefit from ad supported tech are search and video, the rest would be better without ads. Reddit would be a better place if they didn't try to get ad revenue etc, in those cases them chasing revenue makes user experience worse instead of better.
I can't say much about US. As I see it, EU pretty much copied US about that part. There was nothing related to computation in the EU's AI Act projects until few months ago, it was purely a "what kind of data processing are you allowed to do?"
https://www.whitehouse.gov/briefing-room/presidential-action...
"Until such technical conditions are defined, the Secretary shall require compliance with these reporting requirements for:
(i) any model that was trained using a quantity of computing power greater than 1026 integer or floating-point operations, or using primarily biological sequence data and using a quantity of computing power greater than 1023 integer or floating-point operations[...]"
EU:https://thefuturesociety.org/wp-content/uploads/2023/12/EU-A...
> 1023
Should be 10^26 and 10^23.
For example, the parent commenter could have talked about the specific attributes of that model that make it superior. I personally am aware that Mixtral is one of the best performing models right now, but is everyone else? Also, does Mixtral need to be uncensored? I've used vanilla Mistral for some...interesting...prompts and had no issues with it moralizing at me.
This whole thing is a fork of llamacpp, also hoping it'll all go upstream sooner or later.
It's true you can also build dual A6000 with 48+48 = 96GB VRAM also, but that's $10k+ setup just for GPUs on legacy generation.
Runs pretty good on most consumer-grade GPUs, but so far it only supports Windows OS.
On my AMD Ryzen 5 5600u with dual-channel DDR4, I’m getting 2 tokens/second. My friend with Intel Core i3 and single-channel memory was getting 1 token/second.
For all the love llama.cpp gets, its method of dGPU offloading (prompt processing on GPU and then just splitting the model down the middle) is relatively simple. But its interesting that there even is so much "activation sparsity" to take advantage of. The traditional thinking in ML is that memory access is very random.
Hopefully the "cold" neurons eventually get offloaded to the IGP instead?
Also, its curious that they are considering a Metal kernel. I thought the performance advantage came from the hybrid memory pool... seems like that would only help old AMD Macs, unless I am missing something?
A 10x speed improvement is really impressive. If this kind of improvement is reproducible across other models, then presumably identifying hot and cold neurons for inference optimization should go on to become a normal part of model development process.
We have tested PowerInfer on the following platforms:
x86-64 CPU (with AVX2 instructions) on Linux
x86-64 CPU and NVIDIA GPU on Linux
Apple M Chips on macOS (As we do not optimize for Mac, the performance improvement is not significant now.)
And new features coming soon:
Mistral-7B model
Metal backend for sparse inference on macOS
Brilliant!
I see they have a 4-bit benchmark lower down in the page. That's where they ought to compare against exllamav2.
I used pyinstaller. It was difficult because Python makes these things difficult. But it works. It does require an Nvidia GPU. MLC-LLM is another option that might be easier to package and potentially able to run on AMD.
I've been following MLC-LLM as well. Right now I am just using JS/WASM from Huggingface, but later I will want something more performant.
Also apparently exllama has a few side effects in coherence https://www.reddit.com/r/LocalLLaMA/comments/17w57eu/llm_for...
https://github.com/ggerganov/llama.cpp/pull/4543 [Review] Merge PowerInfer with llama.cpp mainline #4543
https://github.com/ggerganov/llama.cpp/discussions/4534#disc... "The x11 speedup is kind of cherrypicked because the llama.cpp GPU code for Falcon 40b is just not well-optimized."
I still think a comparison with exllamav2 or other optimized inference library would make sense too.
Anyway, the technique that they describe is a general one that should broadly improve the ability to run larger models on smaller GPUs by drastically improving perf for CPU offloading. They demonstrate it using both 4090 running the largest models at fp16, but also 2080Ti running the same 4-bit quantized, and they still get ~3x speedup for LLaMA. So this very much sounds like it'll make 33B models the new default on the desktops, while people with even a single 3090 or 4090 will now be able to run 70B at realtime chat speeds.
Does this means that it runs at same time at both CPU and GPU, being faster than a CPU-only or a GPU-only implementation on the same device?
edit: when running on integrated GPUs, can this benefit from the improved communication between CPU and GPU?
But if you want to run a model that requires more VRAM than you have, the current approach is to use llama.cpp and specify n_gpu_layers. That works, but is slower than GPU-only.
OP claims to be 10x as fast as llama.cpp in the case when you can't fit the whole model in VRAM.
Edit: https://socket3.wordpress.com/2016/10/22/using-windows-95-po...