Edit: the above is about PC. Macs are much faster at CPU generation, but not nearly as fast as big GPUs, and their ingestion is still slow.
Recommended reading: Tim Dettmer's guide https://timdettmers.com/2023/01/30/which-gpu-for-deep-learni...
In this video from June, George Hotz says to go with "3090s over 4090s. 3090s have NVLink... 4090s are $1600 and 3090s are 750. RAM bandwidth is about the same." Has he changed his recommendations since then?
It has some upsides in that I can run quantizations larger than 48GB with extended context, or run multiple models at once, but overall I wouldn't strongly recommend it for LLMs over an Intel+2x4090 setup.
It's competitive, but has significant tradeoffs.
> If you factor in electricity costs over a certain time period it might make the Mac even cheaper!
I dunno about that. The M2 Max will happily pull over 200w in GPU-heavy tasks, if we're comparing a 40-series card with CUDA optimizations to Pytorch with Metal Performance Shaders, my performance-per-watt money is on Nvidia's hardware.
If we are talking quantized, I am currently running LLaMA v1 30B at 4 bits on a MacBook Air 24GB ram, which is only a little bit more expensive than what a 24GB 4090 retails for. The 4090 would crush the MacBook Air in tokens/sec, I am sure. It is however completely usable on my MacBook (4 tokens/second, IIRC? I might be off on that).
A 4 bit 70B model should take about 36GB-40GB of RAM so a 64GB MacStudio might still be price competitive with a dual 4090 or 4090 / 3090 split setup. The cheapest Studio with 64GB of RAM is 2,399.00 (USD).
Related, could someone please point me in the right direction on how to run Wizard Vicuna Uncensored or Llama2 13B locally in Linux? I've been searching for a guide and have not found what I need for a beginner like myself. In the Github I referenced the download is only for Mac at the time. I have a Macbook Pro M1 I can use though it's running Debian.
Thank you.
There's a complete list of models at https://gist.github.com/mchiang0610/b959e3c189ec1e948e4f6a1f...
We'll have a better way to browse these soon.
Do you have a guide that you followed and could link it to me or was it just from prior knowledge? Also, do you know if I could run the Wizard Vicuna on it? That model isn't listed on the above page.
https://replicate.com/blog/run-llama-locally
I found that guide here on hn.
I run it cpu only with 16 threads but yeah perf is good enough.
BTw my 6gb figure is me.measuring from htop so llama2 is likely less.
You need about a gig of RAM/nvram per billion parameters (plus some headroom for a context window). Lower precision doesn’t really affect quality.
When Ethereum flipped from proof of work to proof of stake, a lot of used high-end cards hit the market.
4 of them in a cheap server would do the trick. Would be a great business model for some cheap colo to stand up a crap-ton of those and rent while servers to everyone here.
In the meantime if you’re interested in a cheap server as described above, post in this thread.
It might or might not be reasonable speeds, but I would reason that it could avoid "sunk cost irony"; if you decide, that any point, Chat-GPT would have sufficed in your task. It's rare, but it can happen.
If you want to take this silly logic further, you can theoretically run any sized model on any computer. You could even attempt this dumb idea on a computer running Windows 95. I don't care how long it would take; if it takes seven and a half million years for 42 tokens, I would still call it a success!
> thousands for RAM
I wonder if your perspective might be a little off - you can get 64GB DDR4 RAM for ~$100, it’s really not a big deal these days.
It’s a big deal on Mac, of course, where 64GB means big kitted out high-end model that costs thousands, but RAM really is that cheap.
So while a lot of us think that you need to splurge in order to get into LLMs, the reality is you don't, not really, and pretty much any computer will run any model, thanks to the efforts of projects like llama.cpp. Even using the disk like you mentioned! That's a thing, too. It's slower, but it's entirely possible.
If you're willing to drop down to the 7B/13B models, you'll need even less RAM (you can run 7B models with less than 8GB of RAM), and they'll run radically faster.
People have been working really hard to make it possible to run all these models on all sorts of different hardware, and I wouldn't be surprised if Llama 3 comes out in much bigger sizes than even the 70B, since hardware isn't as much of a limitation anymore.
I have been using modal and vast. Vast is cheaper. Modal has some a free inclusions of $30 but that is probably $8 to get the same power in Vast. Modal resell AWS/GCP at the moment. GCP direct seems cheap enough. As does Lambda labs.
With vast some machines don’t start, so you just need to bin them and try another. For learning and non private this is acceptable. For serious stuff I think Vast lets you filter for data-centre GPUs. Modal tends to just work and lets you store the model for later more easily.
Overall: just go with vast. You boot it up and run SSH. It is a familiar experience. Very little time needed on RTFM stuff!