To argue that, I propose a "simplified Chinese room". We have an automaton in the Chinese room, which does only pattern recognition, and can learn by matching patterns and structures, i.e. it will give an answer that is similar to what it has seen before. And we want to teach it to answer a problem "recognize whether the given logical formula is satisfiable". So we give it formulas and it tries to pattern match the formulas it already has seen. Unfortunately, there always exists a formula which is matching closely everything that was shown so far, and yet the correct answer is unexpected. So such an automaton will never be able to discover the actual rule.
What I think has to be added are different levels of discourse, or, a discourse about discourse. With people, we can say, "stop joking around", and it's a meta-discourse signal that we want stop discussing imaginary worlds and focus on the real one, for example, which has to be totally self-consistent. I believe an analogy can be made with Futamura projections, we can have not only systems that learn a specific task (like even GPT-3 does), but also systems that learn to produce an agent to do a specific task, from a metalanguage description that encodes certain rules on how to do it (for example, to follow logical consistency). And if the metalanguage is general enough, you get general intelligence, because then you can apply the agent production process to itself.