Personally I'm so disappointed about the state of local AI. Only old models run "decent" but decent is way to slow to be usable.
When people say "local AI is too slow," they usually mean the engine is too slow, not the model. A 4B model at 186 tok/s (MetalRT on M4 Max) feels genuinely responsive for interactive chat. The same model at 87 tok/s (llama.cpp) feels sluggish. Same weights, same quality, 2x the speed, that's a usability cliff.
We think the gap between cloud and on-device inference is a infrastructure problem, not a model problem. That's what we're working on.