The challenge I've had is I tell it to do {thing}, it does {otherThing} after getting distracted. Then I get annoyed (let's not mention how I probably half-assed the instructions and wouldn't expect a senior human to be able to succeed).
For me, it was a CI/CD issue with our two person startup. I bypass CI/CD often because it was built to catch the AIs. Damn thing went chasing rabbits. Later that day I talk to my cofounder, who says: I have to go chase down this very important CI/CD issue!
Turns out both humans and AIs get distracted relatively easily.
I find that if I think of my AIs less like (deterministic) software and more like leading actually employees that I both get better output and curse (a lot) less.
LLMs _were_ trained on human writing, so it makes sense to me that they tend to act human-like... for better and worse. So yeah, they do dumb stuff, and so do people.
I think we'll get insanely close to being able to "trust" them not to do random stuff, but I am not as confident for jailbreaking still.
Seems to be that Jev is a "reflex" system for AI, where current LLMs are higher level thinking. Computers can now flinch!
Do they sometimes make mistakes? Yeah, and I still check their work. But I've also employed humans, and they make mistakes too. The AI is not worse.
The example above of going through a bug backlog and double-checking closed bugs for accuracy is exactly the kind of work that is excellent for an agent. Assign that task to a normal human being and they would hate your guts. The agent won't protest as long as your token budget is there. You can confirm the results if you want.
I also have a strong feeling we've hit a ceiling on the amount of training data needed for LLMs, what they're all (hopefully) realizing is that you need to focus on how the model reasons, and hopefully someone figures out how to stop people from jailbreaking models, and stops them from just blatantly hacking other companies, that part tells me if it ever were marketed as true AGI, we'd be in very serious trouble.