Waiting for the MTP version to pop up on Unsloth. Speculative decoding makes a huge difference.
Been running quantized 3.6 at 110t/s on a cheap 5060Ti and quite happy with it. If 3.8 improves on it, it would be awesome.
Been running quantized 3.6 at 110t/s on a cheap 5060Ti and quite happy with it. If 3.8 improves on it, it would be awesome.
Honestly, if MTP is available, I don't know why it's so much slower here.