This works well but only if you eyeball the tests and edit them a bit in my experience. Otherwise it gets lazy and makes them trivial to pass. Also, you’ve often gotta explicitly tell it not to hardcode test cases in the solution to make them pass.
You can use property based testing for that.
But I've often run into cases where the AI gets into a vicious spiral of worse and worse code when you keep feeding it the test failures.