I so wanted to be able to use Qwen3.8-27B locally on my fully loaded M1 Max 64GB when spider-mario recommended I use MTPLX[1], but I still found the general MLX MTP acceleration to be so poor, I'd rather just use cheap tokens from OpenCode Zen and Go.
It's just too slow of a model. I know, I know there's new hardware, but my business paid like 4-5 grand for this MBP at the time, and I just don't feel the need to pay 7-8 grand to step up to current hardware when token spend is what it is.