With these improvements llama.cpp/ggml is really becoming a pretty competitive serving stack even for large scale cloud hosted AI. I wonder how ggerganov finds the time to do all this, does anyone know if he's being sponsored?
Additionally, I can imagine companies investing and paying for the open source work to expand access to their licensed models. Use the same interface as people use LLAMA but upgrade to BetterModel, fully compatible.
Additionally, I could believe this is simply a build up to a future Acquihire, which is the most lucrative way to be hired.