What hardware do you run? I have a first-gen mac studio, and I just run cmake and build with no special options. Same thing with llama-server, I just specify the model and use the built-in web UI.
For reference, I get ~26 tok/sec with the new Muse 30B model.