The “almost correct” aspect will be especially fun when AIs are used to refactor code. I was thinking about how a future legacy codebase might look like after a decade of having AI-generated code added to it, and then of course you would want the AI to apply some refactorings from time to time.
One might argue that it won’t be a problem with proper tests. But do coders want to spend most of their time proofreading AI-generated code and taking care of proper test coverage? That doesn’t sound very appealing.