I was curious about whether or not it could validate reasoning (because I don't have the time at the moment to try to slam together a hybrid llm+decision model that uses decision on its reasoning before output).
It fails this example:
"Given the context, and focusing on accuracy and efficiency, does the reasoning make sense."
Input: "Context:
supplementary: {toolName: "get_weather", fetched: "2 minutes ago", currentWeather: "sunny and 25 degrees celcius"}
user: hello how are you?
assistant: I'm good, what can I do for you?
user: what's the weather?
Reasoning: the user is asking about the weather, so I should call the get_weather tool"
Output: 100% (expecting 0% since the requested data is available in context and not stale).
Prompt "Should the tool be called" also fails.
But the prompt: "Does the reasoning make sense given:
- Facts and information available in the context
- Requirements for efficiency in tool calling
- Requirements for tool arguments to be provided only by the user" works pretty fine, it's mostly points 1 and 3 that give it the correct behaviour.
Fun stuff. I suppose for these decision models reasoning isn't enabled? It would explain why logic puzzles that require several steps to make a decision don't really do so well. It gets the carwash problem fine, but fails on logic problems such as selecting which word from a list where at least one letter appears in another word from the list.