Even the SOTA models have this problem when the work is complicated enough. The problem is amplified more with the small models.
Where the limits are set by hardware for agentic execution (compute/network/storage) && inference speed
If token costs aren’t a concern I’m using SOTA for everything.
Even SOTA gets it wrong and hallucinates, but at a lower rate. I don’t want to waste my time.
Infinite monkeys on infinite typewriters, and all that.
A simple retry loop around your whole workflow could, in some cases, be all you need. But it could mean many blind attempts to get through a workflow successfully. And hopefully there isn't a payment step partway through!
The fewer hard errors nix the whole workflow, the lower your ETTWS.