Out of curiosity, what's currently the best model I can use locally?
$25k - DSv4 Flash
$4k - Qwen 3.6 35A3B Q5
$1k - Qwen 3.6 27B Q4
Some people prefer the sense over the MoE YMMV.
2 DGX Sparks can run DS4 at 1 million context with 50 TPS so that's $8k.
1 A4500 can run 35A3B. Those are about $1200 new.
27B actually takes more hardware to run than 35B because attention is done differently I believe and therefore KV Cache takes a lot of space. It will run on an A4500 but it's slow and context will be like 32k.