I had an agent implement a feature the other day. It wrote a bunch of tests proving the correctness of the feature.
It turned out that it had implemented the whole thing in a completely backwards way, which not only defeated the purpose (save CPU) but actually made things worse.
All the tests passed, of course.
I've been thinking for a while, now that the cost of writing proofs via AI is so much cheaper, we can finally fulfil Dijkstra's dream of having all software formally verified.
But in this case, it would have just written a formal proof that the backwards code it wrote was correct!