I tried an oMLX quant of your model (suzu89/Swift-Qwen3.8-27b-oQ8-mtp -- not mine) and liked it. It certainly seems to cut down on thinking compared to stock Qwen3.8 27B in my (limited) testing.
A couple of questions:
- Have you tried the peculiar-ragdoll/Qwen-Sharp-Chat-Templates with it? They replace the default chat_template.jinja with one that encourages less thinking.
- When are you releasing Swift-Qwen3.8-Flash-Next? :)