> Whether they apply "reasoning" or "extremely multidimensional synthesis across hundreds of thousands of existing solutions" is a question of semantics.
I see where you're coming from - if it's good enough that we can't distinguish it, then does any difference really matter? I submit it's fundamentally different. This is essentially the Chinese Room thought experiment [1] or a nice similar metaphor with an octopus from Section 4 of this paper [2].
The trouble is not in its ability but human's interpretation it. Humans see an "avocado chair" and think this AI can invent art and concepts. Producing combinations of existing concepts is not "that hard". Even for combinations that have never existed before.
Meanwhile, it's failing at basic tasks: you can find plenty of examples of it failing basic logic, anything with math or arithmetic, a lot of ethics/bias concerns stemming from the training data, etc.
I think when we look forward to an AI that "reasons" this is not what anybody would mean.
Current AIs are bullshit engines. They are very impressive, and probably even useful. They are a milestone. But they are not reasoning in any meaningful way. And if you look at the math behind them there's really no reason to think they would.
So I guess given a methodology that seemingly shouldn't produce reasoning ability, and no evidence that it has so far, sure, maybe scale will magically unlock it. There's always a chance I guess. But it doesn't really seem too sensible.
See Sam Altman's take as well on Twitter here [3]
[1] https://en.wikipedia.org/wiki/Chinese_room
[2] https://aclanthology.org/2020.acl-main.463.pdf
[3] https://twitter.com/sama/status/1601731295792414720