If people use AI for libraries, OSs, and mission critical software, the apparent productivity gains would have to be weighed against the reliability and performance hits that bubble up to the things that are built on them and rely on them.
I think the Bun port is a great example where testing enabled a very successful implementation. (Both the original tests themselves and runtime comparisons to the previous implementation.)
Tests can show you problems, if you can find them, but they cannot show that there are no problems. Property based testing or fuzzing gets your more coverage, and is a good step, but it is still nothing compared to proving things or understanding how something is built and that it is solid. Testing works towards checking for reliability and robustness, but often it's only 10% (?) of the job.
I'm curious - is there any data on the Bun port error rate? I think that would be very indicative of how successful or not the testing is.
But I feel the need to point out - the goalposts for "does AI work" shift daily.