That's even worse, because now the tests are wrong too! I asked it for a list of test cases for a particular function I wanted to write, and every single one of the cases it presented was incorrect. So it was worse than useless, now I had to check its work and then do the work myself anyway.
It can, but you have to proofread those too.
Why not generate an AI that can take any code and decide whether it has any errors or not?
An AI that tells whether or not a program will halt.