Maybe at 1 or 2 bits of quantization! Even the Macs with the most unified RAM are maxxed out with much smaller models than 405b (especially since it's a dense model and not a MOE).
Anything better than that starts at 200k per machine and goes up from there.
Not something you can run at home, but definitely within the budget of most medium sized firms to buy one.
Got 3.5-4 tokens/s, GPU compute was <20% busy (~90W) and the 16 CPU cores / 32 threads were about 50% busy.
If so, then the parent comment’s sentiment holds true…. Exciting stuff.