if you haven't tried Qwen3.8-Flash-Next with halogen, you're missing out: https://github.com/peonist-ai/halogen-flash-server#the-host-...
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.