But if you go beyond what can be tested easily, asking the agent to do real work rather than writing a patch, imagining things to be true is a problem.
Coding could be treated as a low stakes (time & money consequences for retries) closed loop system where most other tasks cannot.
If it screws up booking your flight/hotel room, how does the agent verify this, and even if it verifies.. there is an actual cost to changes/cancellations.
Similar with agentic e-commerce, lots of ability to screw that up and just seems ripe for fraud / being picked off by bad actors.
I can STILL replicate this behavior in Google AI summaries 10% of the time:
"is <SOMEPLANT> ok for cats"
to which it replies: "Yes, <SOMEPLANT LONG SCIENTIFIC NAME VERBOSE PHRASING> is toxic for cats"
The other one going around this weekend: "how long hot dogs on grill"
Summary: "The hot dogs on your grill are likely around 5-6 inches long .. "
So scale this category of error to unsupervised agents with access to your credit card.
If LLMs worked the way people want to believe they do, there’d be no reason to start in the wrong place — a computer should have the facts!
Unfortunately, travel keeps getting less flexible, with worse cancelation policies.
Trains are usually different because they are much cheaper to operate per trip (for the train operator, not the track operators, but that's a different discussion), so running a half empty train is much less of a problem - especially since you don't need to plan for how much fuel to use ahead of time.
Another example - agentic food ordering. How much more convenient would this make your life vs how much of an error rate would you tolerate if the cost/repercussions are on you?
Would a customer be happy if 2% of the time it sends 20 pizzas to a random address in their contacts list instead of 2 pizzas to their own home? Or 5% of the time it completely ignores your dietary restrictions/allergies and orders an entire meal of food you explicitly told it you cannot eat?
Real world problems don't go away just because it would make the tech neater & tidier.
Only with an LLM that's actually at agent-quality.
If "useful chatbot" and "useful agent" are two rungs on a ladder, the rung before them is "useful autocomplete". Autocomplete that only gets the next token right 90% of the time won't give you compiling code.
The elaborate workarounds you have to build to help an agent which fundamentally doesn’t know what it’s doing reminds me of this old blog post about TDD: https://pindancing.blogspot.com/2009/09/sudoku-in-coders-at-...
IMO present technology is tailored for an experienced developer to give agents manageable tasks that can be one-shot. The marketing right now reminds me of the 90s when AskJeeves promised natural language search when the technology was fundamentally still stuck in keyword search, and learning to craft a search query for Google is today’s prompt engineering