It has some upsides in that I can run quantizations larger than 48GB with extended context, or run multiple models at once, but overall I wouldn't strongly recommend it for LLMs over an Intel+2x4090 setup.
It's competitive, but has significant tradeoffs.
> If you factor in electricity costs over a certain time period it might make the Mac even cheaper!
I dunno about that. The M2 Max will happily pull over 200w in GPU-heavy tasks, if we're comparing a 40-series card with CUDA optimizations to Pytorch with Metal Performance Shaders, my performance-per-watt money is on Nvidia's hardware.
If we are talking quantized, I am currently running LLaMA v1 30B at 4 bits on a MacBook Air 24GB ram, which is only a little bit more expensive than what a 24GB 4090 retails for. The 4090 would crush the MacBook Air in tokens/sec, I am sure. It is however completely usable on my MacBook (4 tokens/second, IIRC? I might be off on that).
A 4 bit 70B model should take about 36GB-40GB of RAM so a 64GB MacStudio might still be price competitive with a dual 4090 or 4090 / 3090 split setup. The cheapest Studio with 64GB of RAM is 2,399.00 (USD).