GPU Guide (For AI Use-Cases)
gpus.llm-utils.org
gpus.llm-utils.org
Stable Diffusion runs well on small GPUs... My 6GB laptop 2060 can run reasonably big resolutions. It can run Facebook AITemplate inference or train a LORA. A 12GB-24GB GPU can run pretty much anything with no issues or fuss.
A single 3090 (or 7900 XTX) can run LLaMA 65B reasonably quickly with llama.cpp. 2x 3090s can run it very quickly with exLLaMA.
MPT and Falcon quantization are extremely new (and TBH the finetunes are not as good as LLaMA yet), but MPT 30b will fit on a single 3090 (or a smaller GPU with llama.cpp) and Falcon (I think) can be split with GPTQ.
Based on your comment, I added a note that the GPU recommendations are for GPU clouds (where if you're running stable diffusion, ~$0.50/hr for a 4090 is pretty practical, and it's not worth going to a cheaper card for most), and I added this note re local GPUs now:
If you're using a local GPU:
* Same as above, but you probably won't be able to train or fine tune an LLM!
* Most of the open LLMs have versions available that can run on lower VRAM cards e.g. [Falcon on a 40GB card](https://huggingface.co/TheBloke/falcon-40b-instruct-GPTQ)
* Thanks Bruce for prompting me to add this sectionI think this can be solved by running RTX 4000 cards with CUDA 12.1 and pytorch nightly cu121... But thats a very uncommon setup.
Also, if we are talking cloud deployments, cheap CPU instances with 32GB-64GB of RAM are a somewhat sane way to test LLMs, thanks to llama.cpp. A GPU will give you better throughput, of course.
But CPU support for Falcon and MPT specifically are in flux, see this post for instance: https://huggingface.co/TheBloke/mpt-30B-instruct-GGML
Additionally, prompt processing will work with large models even with low vram GPUs.
Apple Silicon is interesting for AI because the GPU shares main memory, so you can get one with a lot of RAM and use either CPU or GPU. Less support for Apple GPUs though, at least so far. Work is being done. M1/M2 Macs with 64GiB of RAM are cheaper than many of these GPUs.
Also wow does NVIDIA have near monopoly on this space. Prices will probably come down once Intel and AMD get their acts together.
There are some worrying reports about Intel Arc gen 2, and AMD is just matching price/VRAM with Nvidia these days.
Some of the AI startups are really exotic/business scale only (like Cerebras) or MIA in the consumer/cheap cloud space (like Tenstorrent).
...There are some rumors of an M2 Pro Like APU from AMD/Intel? I hate so sound so grim, but thats about the only positive news I got.
https://ark.intel.com/content/www/us/en/ark/products/232592/...
I've seen some interesting work on deploying models to FPGAs, but high-end FPGAs are also expensive.
As long as chatGPT still lets people ask medical/healthcare questions, it doesnt seem to urgent. I'm dreading when the AMA lobbies to ban it unless under Physician supervision.
There are restrictions on NVENC streams with consumer cards, but that has been a solved problem for a while [0].
If they were to make a consumer card with more VRAM, it would immediately undercut their own Quadro/Tesla lineup, which cost substantially more. I don't see a reason for them to do it.
Money
Apple's neural engine is really interesting as a GPU alternative. It's still just a bunch of tensor cores, but the unified memory architecture allows it to use the full system RAM without a major performance hit. I think we will see more of this in the future.
Also would be nice to compare some of the local GPU options (including Macbooks) vs cloud options.
Edit: I just noticed https://gpus.llm-utils.org/recommended-gpus-and-gpu-clouds-f...
Perhaps the writer kept them unmentioned in an attempt to reduce demand?
The drivers were pretty easy to find. I think I just googled them.