Sincere question: do we know enough about human intelligence to support that last assertion?
I would argue that we are not - because we do not have a rigid partition between training/test. Our ability to reason speculatively rests to some extent on ability to produce outputs that we have not seen before and which are statistically not like that which we have already seen. How we know which of these outputs to keep, and which to discard is (afaik) an open question.