Where do you put the 480 GB to run it at any kind of speed? You have that much RAM?
Or you can rent a newer one for $300/mo on the cloud
The RAM is for the 400gb of experts.
For reference, a RTX 3090 has about 900GB/sec memory bandwidth, and a Mac Studio 512GB has 819GB/sec memory bandwidth.
So you just need a workstation with 8 channel DDR5 memory, and 8 sticks of RAM, and stick a 3090 GPU inside of it. Should be cheaper than $5000, for 512GB of DDR5-6400 that runs at a memory bandwidth of 409GB/sec, plus a RTX 3090.
Qwen says this is similar in coding performance to Sonnet 4, not Opus.
If you have 500GB of SSD, llama.cpp does disk offloading -> it'll be slow though less than 1 token / s
3 t/s isn't going to be a lot of fun to use.