What do you use instead?
https://github.com/ggml-org/llama.cpp https://github.com/mostlygeek/llama-swap
https://omlx.ai https://vmlx.net
That said, it's been a few weeks since I've looked so maybe llama.cpp has those features now... they really do move that quickly.