Sounds interesting would love to test it out but here's what I got when I first tried to get some models downloaded.
On a 64gb Ram and 5080 machine the biggest model suggested was Qwen 9B
can hit 90+ tps on the MoE 35b but mag thinks it wont fit.
On a 64gb Ram and 5080 machine the biggest model suggested was Qwen 9B
can hit 90+ tps on the MoE 35b but mag thinks it wont fit.
The MoE 35b might be tight though unless you were to go below 4-bit. Could you share the quant you used when you ran this on that 5080 before? Our catalog only contains models down to 4-bit because we find that thinking, tool calling, and overall capabilities start to suffer at lower fidelity.
Feel free also to create a GitHub issue with more details and we can take a closer look.