ML systems do not guarantee valid output. If you have a system that can prove a given transformation is behavior-preserving, then sure you can let the AI have a go at it and then check the result, but having such a checker is much harder than it sounds.