It fails this example: "Given the context, and focusing on accuracy and efficiency, does the reasoning make sense."
Input: "Context: supplementary: {toolName: "get_weather", fetched: "2 minutes ago", currentWeather: "sunny and 25 degrees celcius"} user: hello how are you? assistant: I'm good, what can I do for you? user: what's the weather? Reasoning: the user is asking about the weather, so I should call the get_weather tool"
Output: 100% (expecting 0% since the requested data is available in context and not stale).
Prompt "Should the tool be called" also fails.
But the prompt: "Does the reasoning make sense given: - Facts and information available in the context - Requirements for efficiency in tool calling - Requirements for tool arguments to be provided only by the user" works pretty fine, it's mostly points 1 and 3 that give it the correct behaviour.
Fun stuff. I suppose for these decision models reasoning isn't enabled? It would explain why logic puzzles that require several steps to make a decision don't really do so well. It gets the carwash problem fine, but fails on logic problems such as selecting which word from a list where at least one letter appears in another word from the list.