Says who? It’s quite easy to limit the blast radius of a reasoning error.
Says who? It’s quite easy to limit the blast radius of a reasoning error.
Sure, you can patch that specific case with guardrails, but how many unpredictable edge cases are you going to cover? It only takes a user with a bit of ingenuity to circumvent them. There are already several examples of AI agents getting stuck in infinite loops, burning through massive API bills while achieving absolutely nothing.
You can contain a system failure, but you cannot contain a logic failure if the system doesn't know the logic is wrong.
It didn't happen. Seems the bug was "contained".
Sort of undermines your point re "catastrophic business failure" don't you think?
This is the wrong question. The correct question is what specific subsets of cases do you allow, similar to any security question
Suppose you had:
Math() Add() Subtract()
Program() Math(“calculate rate”)
This is intentionally written vaguely. How do you limit that these implementations ensure Program() runs and does the right thing when there is no guarantee Math() or its components are correct?
Normally you could use a typed programming language, unit tests, etc, but if LLM is the ultimate abstraction programs will be written line above. At some point traditional software engineering principles will need to apply.