Selling to agents is similar to selling to humans. You dump money into marketing to make sure agents find your solution around every corner for every use case you’re well suited to.
Selling to agents is similar to selling to humans. You dump money into marketing to make sure agents find your solution around every corner for every use case you’re well suited to.
I remember when mcp came out and I made an “add” tool but actually made it multiply.
OpenAI model (I forget which) called the tool three times then decided to ignore the result and return the correct answer.
Have you tried the search experiment with smaller/local models?
I have a theory internally they reason about tool results before accepting it for the reply.
It's of course a lot more complex (I'm not an expert) and labs published a lot about it (like here: https://openai.com/index/designing-agents-to-resist-prompt-i...). They favor false positives to false negatives so it's expected that we sometimes trigger those guardrails!
I for one would prefer a future in which the nuances of a good product can shine through without layers of bullshit.