If I can use it with Linux in any meaningful way, that would be a better selling point.
I posted it here the same day I found and started using it, to almost no reaction.
[0] https://github.com/FastFlowLM https://fastflowlm.com/ https://huggingface.co/FastFlowLM
HN is overloaded with AI stuff, its hard to break through all the noise. I say this as someone very interested in AI. Even I skip some links because its just too much.
some older pre-395 AMD articles suggested it'd be possible to use the NPU for prefill and the GPU for decoding and this would be faster than using either alone, but we have yet to see that (even on Windows) for any usefully sized models, just toys like LLaMA-8B.