A human does not do this.
First of all, most questions we have been asked before. We have made mistakes in answering them before, and we remember these, so we don’t repeat them.
Secondly, we (at least some of us) think before we speak. We have an initial reaction to the question, and before expressing it, we relate that thought to other things we know. We may do “sanity checks“ internally, often habitually without even realizing it.
Therefore, we should not expect an LLM to generate the correct answer immediately without giving it space for reflection.
In fact, if you observe your thinking, you might notice that your thought process often takes on different roles and personas. Rarely do you answer a question from just one persona. Instead, most of your answers are the result of internal discussion and compromise.
We also create additional context, such as imagining the consequences of saying the answer we have in mind. Thoughts like that are only possible once an initial “draft” answer is formed in your head.
So, to evaluate the intelligence of an LLM based on its first “gut reaction” to a prompt is probably misguided.
Let me know if you need any further revisions!