[flagged]
For reference, I get ~26 tok/sec with the new Muse 30B model.
In the past, I had to play with chat templates for some models to work with agents for tool calling. But I've never had to do anything other than specify the model, and tweaking the context size in some cases.