How to train large models on many GPUs? (2021)
lilianweng.github.io
lilianweng.github.io
FYI, I am not affiliated with Ray. However, I did write the following paper on scaling data-parallel training for large ML models ;) https://openreview.net/pdf?id=rygFWAEFwS
Also, another one of my papers talks about distributed training while reducing the communication bottleneck for distributed training: https://dl.acm.org/doi/pdf/10.1145/3447548.3467080
If you're curious about how Ray is used for LLMs, here are some interesting examples of LLM projects using Ray!
- Alpa does training and serving with 175B parameter models https://github.com/alpa-projects/alpa
- GPT-J https://github.com/kingoflolz/mesh-transformer-jax
- Another HN thread on training LLMs with Ray (on TPUs in this case) https://news.ycombinator.com/item?id=27731168
- OpenAI fireside chat on the evolution of their infrastructure and usage of Ray for training https://www.youtube.com/watch?v=CqiL5QQnN64
- Cohere on their architecture for training LLMs https://www.youtube.com/watch?v=For8yLkZP5w&t=3s
Some other thoughts
1. There is a lot more we want to do to make Ray better for working with large language models and for making training, serving, and batch inference work well out of the box.
2. The original post is about training, but we actually see even more interest in fine-tuning and serving with LLMs, in part because there are good pre-trained models.
3. For LLMs, we see a lot of interest in Ray + Jax or Ray + TPUs relative to what we see in other use cases.
I tried torch FSDP but it only managed to increase the memory to something like 150% of 1 GPU.
I eventually ended up sharding my model manually with .cuda() and .to() which works much better, but now I am limited to one module on one GPU and I would like to expand even more, and that would mean spinning up more nodes and splitting the model over how many GPUs manually.
I would be interested if anyone knows of a framework that manages this automatically and just works.
EDIT: BTW I am talking about model sharding not data parallelism which works very well with DDP.
Nvidia's NCCL and AMD's RCCL provide parallelism constructs that really are hidden at the framework level (such as PyT).
However, I don't think that you would want to hide model, data, or tensor parallelism. It's too important a consideration for performance and training convergence impact.
At least in scientific computing, I've never observed effective means of automatic parallelism expressed across many nodes despite decades of research. I'm not optimistic this will be effective anytime soon.
https://huggingface.co/transformers/v4.9.2/parallelism.html#...
Check MosaicML if it might help in your case. I haven’t tried myself but they’ve most customizations and speed up optimizations I came across in the recent times
https://www.mosaicml.com/blog/supercharge-training-composer
Also worth checking out their “training from scratch” blog posts.
Training StableDiffusion: https://www.mosaicml.com/blog/training-stable-diffusion-from...
Training GPT-3: https://www.mosaicml.com/blog/billion-parameter-gpt-training...
* It gives you PyTorch DDP for free. Makes FSDP about as easy as can be, and provides best in class performance monitoring tools. https://docs.mosaicml.com/en/v0.12.1/notes/distributed_train...
Here's a nice intro to using Huggingface models: https://docs.mosaicml.com/en/v0.12.1/examples/finetune_huggi...
I'm just a huge fan of their developer experience. It's up there with Transformers and Datasets as the nicest tools to use.
Are these specific libraries?
Supports data, tensor, pipeline, sequence parallelisms, activation checkpointing, distributed optimizers, fused kernels and more.
Question: could this be implemented in PyTorch in an opaque way? Or would it require changes to its API?
How to train large models on many GPUs? - https://news.ycombinator.com/item?id=28657797 - Sept 2021 (9 comments)
Her follow up post [1] is also recommended for those who (like me, are experienced but not in ML) finally had things click because of the OP writeup:
Large Transformer Model Inference Optimization (2023)
https://lilianweng.github.io/posts/2023-01-10-inference-opti...
A very cool cite from that article is LLM.int8(): https://arxiv.org/abs/2208.07339
https://www.youtube.com/watch?v=-fEjoJO4lEM
In principle they could be used with an API like DirectStorage RDMA or CUDA GPUDirect RDMA (which dates back to Kepler) and in this case they would never need to talk to the CPU, given appropriate software support. But it's not going to be presented as GPU memory ever, it's going to work like a block storage device you can do RDMA requests against, most likely.
https://docs.nvidia.com/cuda/gpudirect-rdma/
https://developer.download.nvidia.com/video/gputechconf/gtc/...
Now technically - it all depends on what you mean by "as GPU memory" because PCIe is all RDMA anyway, even CPU-to-GPU is a RDMA operation. That's why there's the whole thing about "resizable BAR" etc - that's the aperture window in CPU memory that gets mapped in from the GPU memory.
So technically yes you can map those SSDs in "as GPU memory" via GPUDirect RDMA block storage (or DirectStorage), but you can do that with a regular NVMe SSD in an adapter card too. The SSG is just a "combo GPU+SSD card" in the same way QNAP makes those "combo network+SSD cards", but with a lot of fanfare/marketing around it.
To be absolutely fair, Fiji/Vega is a good design for that since it doesn't have a bunch of memory packages around it. But it wasn't what AMD trumpeted it as, as the LTT video describes, it was a very specific reaction to the question of 'workstation GPUs are using a lot of memory, HBM can't be scaled as high, how do we put more memory on a Fiji/Vega GPU for workstation users". And AMD's claims that HBM meant you could just swap everything around and not have to worry about framebuffer size were never true, the PCIe bus itself is not fast enough for that.
So if you are building GPUs or AI accelerators, you tend to just go ahead and build this in.
https://www.servethehome.com/wp-content/uploads/2022/08/NVID...
That unless an utterly revolutionary new interconnect technology comes...
LPDDR is getting adopted more and more too, with CPUs losing memory expansion capabilities in exchange of huge power savings.
Or just a plain slower GPU: Apple's GPU series, but that's at very much higher price tags than desktop GPUs at a given perf level.