Since I couldn't find it in your list, I'd like to plug my own macOS (and iOS) app: Private LLM. Unlike almost every other app in the space, it isn't based on llama.cpp (we use mlc-llm) or naive RTN quantized models (we use OmniQuant). Also, the app has deep integrations with macOS and iOS (Shortcuts, Siri, macOS Services, etc).
Incidentally, it currently runs Mixtral 8x7B Instruct[2] and Mistral[3] models faster than any other macOS app. The comparison videos are with Ollama, but it generalizes well to almost every other macOS app that I've seen uses llama.cpp for inference. :)
nb: Mixtral 8x7B Instruct requires an Apple Silicon Mac with at least 32GB of RAM.
[1]: https://privatellm.app/
[2]: https://www.youtube.com/watch?v=CdbxM3rkxtc
[3]: https://www.youtube.com/watch?v=UIKOjE9NJU4