For training, yeah. But at least in some cases, inference can be optimized to use less memory.
For their large Whisper model, OpenAI says they need 10GB of VRAM: https://github.com/openai/whisper#available-models-and-langu...
My DirectCompute port of that uses slightly over 4GB of VRAM, see the last column in that table: https://github.com/Const-me/Whisper/blob/master/SampleClips/... I haven’t actually optimized memory usage I only did the DirectCompute port, the memory savings were achieved by Georgi Gerganov in whisper.cpp project, which I ported to Windows.
https://nlpcloud.com/gpt-3-open-source-alternatives-gpt-j-gp...
though TBH just making cards with more RAM to fit these kind of models seems easier than being at absolute cutting edge of performance
I also wonder how this metric translates to something like Apple Silicon 'unified memory'... latest MBP have 32GB RAM, I wonder if anyone managed to run GPT-J there yet?
https://github.com/ggerganov/ggml/tree/master/examples/gpt-j
it's CPU-only and runs inference at "about ~6 words per second" on an M1 MBP
These models are obviously nerfed, but you can run them on budget ARM hardware and get legible, paragraph-length responses in less than 10 seconds.
Thank you for your work.
For the A100, A6000s, and H100s you are paying for the memory and the upcharge for being a server card. AMD isn't really competing in this space, but they are moving in that direction if you look at the exascale machines[0,1]
[0] https://www.top500.org/system/180047/
[1] https://www.amd.com/en/products/server-accelerators/instinct...
How much time is saved?
If you've got a 12GB card at 5x speed vs. a 96GB card at 1x speed (so 5x slower), is the former still competitive? How much faster does the 12GB card need to be to get the same quality results with less memory (and smaller batches)?
What if the faster card has 24GB of memory?
In classification batch is going to help a ton (including stability) and you may not even get top results without a high batch, so infinite in that case. More realistically, consider that batch also helps with GPU I/O. I/O slows things down quite a lot, so if you're moving stuff in and out of the accelerator a lot then you can easily be I/O bound.
For generation the conversation gets more complicated. For GANs the smaller card could do better if you can fit a proper batch size there (probably not tbh but often closer to 16 or 24 so that may be better here). Here distributed tends to help more. For VAEs, Diffusion, Autoregressive, etc are more stable and can also hugely benefit from large batches.
So the truth is an unsatisfying "it depends." I know that may not be the answer you're looking for, but it is the real one. It takes a lot of work and tricks to scale (as anyone working in non-ML HPC will tell you very similar stories about scaling). But also note that the clock speeds aren't as big of a gap as you note though the memory gaps are.
a larger Batch size can help keep the gpu fed but usually that’s not a problem with these larger models.
Between the two mentioned cards, though, it really does come down to performance. The cards are simply not interchangeable for large scale model training.
This software is mostly supplied by vendor, and it mostly cares about vendor's tech, and will usually require special features exposed by the hardware, so, won't work with consumer-grade GPUs.