How do you quantify capability though? In coding or math there's an easy way to validate the output. That makes it easy to judge capability. But for many cognitive tasks the output is either subjective, not quantifiable, or requires deterministic results. No validation can tell you if an essay is good. Or if an argument will be persuasive to a specific audience. Or that the statistics an LLM pulled from a data source are accurate.
So how does an LLM learn to outperform humans when its output in a large number of tasks can't be validated? These sorts of tasks are a large part of cognitive work and intelligence to me.