I'm also looking into expanding the protocol and the engine to support various steering techniques.
I'm also looking into expanding the protocol and the engine to support various steering techniques.
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself.
As I said, I don't have the hardware to test it myself, so let me know how it goes!
I've submitted it as PR, as well as an initial implementation for an OpenAI like API
Maybe Intel and AMD should help them with that.
I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?