Asked it to spot check a simple rate limiter I wrote in TS. Super basic algorithm: let one action through every 250ms at least, sleeping if necessary. It found bogus errors in my code 3 times because it failed to see that I was using a mutex to prevent reentrancy. This was about 12 lines of code in total.
My rubber duck debugging session was insightful only because I had to reason through the lack of understanding on its part and argue with it.
Try again with Sonnet 4
Try again with GPT-4.1
Here I thought these things were supposed to be able to handle twelve lines of code, but they just get worse.
But which codebase is perfect, really?
I should probably stop commenting on AI posts because when I try to help others get the most out of agents I usually just get down voted like now. People want to hate on AI, not learn how to use it.
And it requires a bit of prompt engineering like using caps for some stuff (ALWAYS), etc.
GPT-5.2-Codex did a bad job of obeying my more detailed AGENTS.md files but GPT-5.3-Codex very evidently follows it well.
I find it infinitely frustrating to attempt to make these piece of shit “agents” do basic things like running the unit/integrations tests after making changes.
We’ve been acting as if it’s assembly code that the agents execute without question or confusion, but it’s just some more text.
people still do useful work without a global view, and there's still a human in the loop witth the same ole amount of global view as they ever had.