So figure out how to run it on Vulkan instead of requiring the user to be locked into expensive CUDA cards.
So figure out how to run it on Vulkan instead of requiring the user to be locked into expensive CUDA cards.
This repo is a vibecoded demo implementation of some recent research papers combined with some optimizations that sacrifice quality for speed to get a big number that looks impressive. The 207 tok/s number they're claiming only appears in the headline. The results they show are half that or less, so I already don't trust anything they're saying they accomplished.
If you want to run Qwen3.5-27B you can do it with a project llama.cpp on CUDA, Vulkan, Apple, or even CPU.
I just find it funny they talk about being vendor locked, and the only thing they support is nvidia.
Anyway, I don't have any Nvidia hardware, and I've got several local models running and/or training at all times.
It's just funny they talk about vendor lock, and they only support nvidia.
The built in Apple Intelligence right now is very small, but even just having a small LLM you know is always there, online, fast and ready makes you think about building app differently. I would love the context to expand from the meager ~4K tokens.
Metal is Apple's API.