At work I am sitting on a small pile of incomplete/wrong/missing-the-point bug reports right now, all generated by Opus 5 on High effort. Even under good conditions, LLMs are still wrong quite a lot, and confidently so. I can see why you think instant dismissal of LLM generated work is shallow, but I think it's at least as shortsighted to assume that when it creates poor quality work it must be an old/cheap model, bad settings, bad prompting, etc.