In text mode I'd use structured output for this, or even get away with instructing the model to [emote] or even just map emojis to emotes.
Great demo though, I'm going to try out the image stuff now. Super rad that it's multimodal decision making; "does this pipe need to be inspected?"->"yes","I want a closer look","no" etc!
Edit: impressive stuff! I gave it a bunch of examples: - things that humans should/should not eat and it got this correct (including rejecting rat poison) - probability staff member should be called to a train platform (people standing vs. child looking over the edge)
It fails this example: "Given the context, and focusing on accuracy and efficiency, does the reasoning make sense."
Input: "Context: supplementary: {toolName: "get_weather", fetched: "2 minutes ago", currentWeather: "sunny and 25 degrees celcius"} user: hello how are you? assistant: I'm good, what can I do for you? user: what's the weather? Reasoning: the user is asking about the weather, so I should call the get_weather tool"
Output: 100% (expecting 0% since the requested data is available in context and not stale).
Prompt "Should the tool be called" also fails.
But the prompt: "Does the reasoning make sense given: - Facts and information available in the context - Requirements for efficiency in tool calling - Requirements for tool arguments to be provided only by the user" works pretty fine, it's mostly points 1 and 3 that give it the correct behaviour.
Fun stuff. I suppose for these decision models reasoning isn't enabled? It would explain why logic puzzles that require several steps to make a decision don't really do so well. It gets the carwash problem fine, but fails on logic problems such as selecting which word from a list where at least one letter appears in another word from the list.