On a Mac the baseline I'd want is MLX, not llama.cpp. llama.cpp isn't the fast path on Apple Silicon for most models people run locally, so a speedup over llama.cpp could still be slower than mlx_lm. Do you have that number?
Second, more important for agents: decode speed is rarely what hurts. It's resending the same system prompt plus tool schemas every turn. Does self-optimizing cover prefix cache reuse across requests, or is it kernel and layout tuning only?
And is the endpoint OpenAI-compatible? My runtime already talks to llama.cpp and LM Studio through that wrapper, so drop-in is the difference between trying it tonight and not.