30B/q4 requires 20GB of RAM while 3090 has 24GB.
30B/q4 requires 20GB of RAM while 3090 has 24GB.
Llama 30B 4-bit has amazing performance, comparable to GPT-3 quality for my search and novel generating use-cases, and fits on a single 3090. In tandem with 3rd party applications such as Llama Index and the Alpaca LoRa, GPT-3 (and potentially GPT-4) has already been democratized in my eyes.
Need 64GB for 30B
And 128GB for the 65B
I'm not sure about the performance, but I think it should be ok? Especially given how much Apple has been investing in what -- I believe they call -- neural cores?
Here's some more context: https://news.ycombinator.com/item?id=35105364
What benefit does that offer over 2 GPUs?
Meanwhile the A6000 VRAM bandwidth is 768 GB/s (= 16 Gb/s (GDDR6) × 384 bit-width ÷ 8 bits per byte).
(Also, the RTX 3090 has faster VRAM, >900 GB/s, than the A6000, because it is GDDR6X.)