Run 70B LLM Inference on a Single 4GB GPU with This New Technique
ai.gopubby.com
ai.gopubby.com
“I try it on a GTX 1060 6Gb on windows, but I don’t think it’s supposed to be that slow? It took me 13 hours to generate one sentence, RAM was around 9.5Gb and only 6% of GPU in use. It works but I don’t know what I’m doing wrong”
Which also has a name that they even use, model sharding. Which is well known to be very slow because you increase the number of IO operations to the GPU. I don't even understand the post because they mention sharding is slow and then say they do sharding but as if they don't. Even the comment in their code calls it sharding. Unless I'm missing something, it's pretty fucking deceptive to call this a new technique. Can we stop with the AI hype? The field is exciting enough that we don't need to be selling snakeoil and selling everything as if it is 10x better than it is.
For anyone interested in model sharding you should look into pytorch's gradient checkpointing documentation and/or fairseq.
Without getting into a bunch of stuff about FDSP or Accelerate or even the NCCL/All-Reduce type ops (which I obviously hope the interested check out) I’ll signal boost my favorite tech talk ever by maybe the coolest CS prof living:
https://youtu.be/l5JqUvTdZts?si=9HlNWHTlRbiatHD1
I’m biased because I’m also Ben, but I’m not “Ben Rekt” who is a musician and charismatic and way too cool to care about SV.
Hogwild is a really cool routine. There's a lot of work going into FDSP, parallelism, and all kinds of things at the low level for optimizing communication, maximizing parallelization, writing cuda kernels, and so on. I think there's a lot of unsung heros of ML working in these areas that people never even see (including researchers). So many people quote Sutton's Bitter Lesson but these are the people that made that landscape even possible. Without these people working at the low level we wouldn't be able to process large models or large datasets.
The fame issue is always a bit weird to me. Like I think more people know of Vaswani (AIAYN) or Dosovitskiy (ViT paper) than those that know Kingma, who wrote the Adam optimizer. Both these works use Adam as well as Rombach's LDM (Stable Diffusion). I'm only mentioning Kingma because I'm more aware of that work but am certain there's tons of work that could be better highlighted (actually illustrating this point).
If the model is all in RAM and being moved one layer at a time to a GPU, would that be faster than getting it to the CPU? What about if the model is on an SSD. The point being, compute is not the bottleneck so I wonder if this provides speedup compared to an optimized model run on a CPU?
1. The GPU was still faster, though the difference was smaller than I really expected.
2. Almost all the time was spent on compute, not loading weights to/from the GPU. For models I could fit on my GPU the difference in performance in loading one layer at a time vs loading them all was completely negligible, and the difference in memory usage (constraining what else I could do with the computer simultaneously) significant.
3. Loading from SSD was a huge bottleneck when I tried that
Hardware: RTX 2070 super, 5800x
The way to get a speed up is to split the load across both the CPU and GPU since LLMs can be split into layers. This also increases the size the model can be.
Even for non-batched, you can still optimize this a bunch for fairly big models without much more work. Basically, overlap io & compute: while one layer is computing, already be loading in the next layer(s) in parallel. PCI speeds are in the 1GB/s-32GB/s each way for consumer cards nowadays.
Without a lot of work, they might be able to fully hide the IO latency.. very model + hw dependent.
(ofc this assumes that those capabilities considered "emergent" at larger sizes do not require later layers which they probably do). There is a weird situation where we have the sentence transformer, the bge, e5 and etc kind of embeddings, and then a big jump up to generative model embeddings the model providers provide, but not much widespread adoption of e.g GPT-J or neo-20B embeddings (even though at time of release, they had notable usecases over sentence transformers)
One of the benefits of lack of contextual or historical understanding is the lack of barriers to action. One of the downsides is you’re likely to do a lot of things of questionable value.
This is hysterical, they're literally swapping from VRAM to RAM and certainly to disk as well since a 70B GGUF is at least 30-40 GB in total.
I suppose it's slightly better than not being able to to it at all, but jesus christ this is just hilariously bad.
1.3b is about the max you can get for 4gb vram
7b will work on 8gb vram if you quantize
13b will work on 16gb vram if you quantize
33b can work on 24gb if you quanitze
If you lack vram but have ram and infinite patience you can do offloading shenanigans with cpu but it tanks performance to unusable levels in my experience.
I use Phind_34B for a 24GB gpu. When in use I typically set gpu memory limit to 22.5GB.
Quantization: 1hr 3min. 4bit, group_size=32, desc=true, damp=0.01, trust_remote = false, dtype torch.float16, use_triton = false.
Thebloke, turboderp, lonestriker, bartowski. Between these 4 users on huggingface you can find pretty much every common model already quantized.
Uhh... From the abstract:
> We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. —Vaswani, et al. 2017
This is from much earlier this year and seems to be the same idea, only loading part of the model into the GPU at a time.
>"During inference, layers are executed sequentially. The output of the previous layer is the input to the next. Only one layer executes at a time.
Therefore, it is completely unnecessary to keep all layers in GPU memory. We can load whichever layer is needed from disk..."
...Or from the RAM of the main computer the Graphics Card is attached to, via the PCIe Bus and DMA (or whatever other methods exist now or in the future)...
>"...when executing that layer, do all the calculations, and then completely free the memory after."
...Or keep it in the adjacent RAM of the main computer, for more speed...
(Random idea here for hardware designers incidentally:
1) Graphics Card/GPU manufacturers -- put a high-speed CXL (Compute Express Link) interface on your Graphics card (https://en.wikipedia.org/wiki/Compute_Express_Link).
2) Independent Hardware Designers -- design a stand-alone RAM board (or RAM board rack/tower) which has a high-speed CXL interface on it, which can be interfaced to #1, above.
If #1 and #2 exist, now we have Graphics Cards/GPU's that can be interfaced to very large high-speed banks of external RAM -- effectively making the limited local memory of a GPU, a thing of the past...
Yes, there might be a slight latency, slightly above that of local GPU memory when accessing this adjacent memory. But I'm guessing that if this latency exists, then over time this will be engineered away...
In the short term, well designed software and/or caching strategies could compensate for any latency overhead, if there is one...
In conclusion, Graphics Card / GPU + CXL + External Rack o' RAM -- could be a very winning combination...)
>"This way, the GPU memory required per layer is only about the parameter size of one transformer layer, 1/80 of the full model, around 1.6GB."
A brilliant observation and a brilliantly written article!
Well done!
https://developer.nvidia.com/blog/simplifying-gpu-applicatio...
This would allow you to mmap() the weights file into CPU memory, and pass that CPU pointer to the GPU and allow it to fault the data in on-demand.
The GPU will have less bandwidth accessing this memory than if you copied it onto the GPU though, as it will need to take a trip over PCIe to read it. But it would eliminate the need to manually upload data to the GPU and the synchronisation involved with that.
For GPUs to take data from CPU RAM, it must traverse PCIe which is on the order of ~100x slower than through the on board VRAM.
I don’t think the GPU can address all the memory though, I know that the 128GB is limited to 96GB that can be shared with the GPU.
The M2 Ultra in the Mac Studio has 800 GB/s. The Nvidia A40 has 696 GB/s while the Nvidia H100 SMX has 3.35 TB/s, which is probably the best you could get today.
https://www.intel.com/content/www/us/en/support/articles/000...
Or you could get a Mac Studio for 4 800 USD.
Nvidia supports reading data directly from main memory. They introduced it around 2013 as Unified Memory, and it occasionally gets updates that make it more usable. It's probably faster than loading each layer into VRAM, computing it, discarding the weights and loading the next layer; but it's still a lot slower than having the weights in VRAM and often slower than just doing the work on the CPU.
Alright, nothing to see here.
I could really use some kind of filter for stuff that does work with a beefy consumer GPU and stuff that does not.
You don't need 16GB of VRAM for this, it's just what they had available for testing.
Btw, has anyone experimented with Microsoft's DirectStorage api for this usecase? I understand that the shards will not be very compressable, but maybe the increased IO offloading can have benefits?
But anyway, if someone has $80B I can get this done. It's just an optimization problem.
* Yes, performance numbers would be helpful.
* No, this is not new, and for most applications, not practical.
The basic point is that yesterday, I couldn't run a large model locally. Today, I can, in about one pip install and 20 lines of code. Yay!
There are a lot of things which are convenient to be able to test locally. YES it takes overnight, but if you're trying to confirm if there's a business case, that's often good enough. For a single compute (or even just a few), this is also simpler than spinning up a cloud machine.
And you can do it in MLC, in your IGP, if you have enough CPU RAM to fit the model.
Running a 70B very slowly is nothing new. To be blunt, this strategy is a bad idea in the face of newer implementations.