I wonder if the issue is that not many people within OpenAI can validate the output. Therefore what reads well looks good, and then was published.
If that’s the case in a way its a similar delusion that average people are experiencing with their own AI use.