This is a fancy way of saying that “the problem is tool calling”, which is obviously true. The problem is that, when it works correctly (99.99% of the time), it adds so much more value to LLMs.
Sandboxing is a step in the right direction, but can also add friction.
Using guardrails is also good, but adds latency, expenses, and also doesn’t solve 100% of the issues.
IMHO there currently does not exist a proper solution to this problem, and it has yet to be discovered. The proper solution, however, should NOT be based on LLMs, so guardrails are the incorrect direction (albeit effective and easier to implement).