Why mostly unusable 125b?
I assume you are talking about qwen3.8-flash-next. Support for it on some places, like llama.cpp, is still wip (depending on configuration) but it looks like a very capable model in it's category.
I assume you are talking about qwen3.8-flash-next. Support for it on some places, like llama.cpp, is still wip (depending on configuration) but it looks like a very capable model in it's category.
So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people.
With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow.
Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).