Custom tool calling formats are iffy in my experience. The models are all reinforcement learned to follow specific ones, so it’s always a battle and feels to me like using the tool wrong.
Have you had good results with the other frontier models?
Have you had good results with the other frontier models?
I’ve tested local models from Qwen, GLM, and Devstral families.
GPT models can follow tool format correctly but don't keep on going.
Grok-4+ are decent but with issues in longer chats.
Kimi 2.5 has issues with it reverting to its RL tool format.