The way to use AI is to make sure it has clear, verifiable success criteria, test suites, etc. Make sure any output has citations, reduce the need for trust to zero, etc.
I see people one shot stuff and it makes no sense, is completely fake half the time, just like you point out.
It should be the case that Codex and Claude Code should incorporate this kind of thing automatically at some point.