Judging these systems solely by their output is to repeat the msitakes of behaviorism, even a dumb markov chain or a parrot converses better than an infant, but unlike the infant does not acquire an understanding or representation of language.
Non-snarky question: What else can you judge by? Isn't any alternative just putting more precise conditions on the output?
With ChatGPT, it's still easy enough to see it's mistakes, and its attempts at fiction and poetry, impressive as they are, are still clumsy to a trained eye, relative to expert human work. But what if they weren't? What happens when they're indistinguishable?
its architecure. A child is a living and autonomous agent. It has (or develops) meta cognition, an awareness of its own mental state (and by extension use of language). These models don't have the capacity to do this even in theory given that they're static and pretrained. When you ask ChatGPT what it feels like to speak, there isn't some neural activity within the model, it has no model of itself that it actively inspects, it doesn't learn while it converses with you, it just tells you what someone wrote on Quora two years ago.
>What happens when they're indistinguishable?
Then the system is likely going to look a lot different than it does now because these aspects of cognition seem pretty important when you want something that is genuinely human-like rather than just mimicry or memorization.
> Then the system is likely going to look a lot different than it does now
I think I agree with this too, but to challenge the idea: I would never have thought "a fancy autocomplete architecture" could give rise to something as sophisticated as ChatGPT, its flaws notwithstanding. So I don't feel so confident that further iterations of the idea, or iterations that involve other architectures that are "still obviously fake" won't give rise to results that far more terrifyingly convincing than ChatGPT.
Since we don't really understand these architectures, human or machine, I don't see how that can be used as the criteria. Ever more find-grained versions of "output" seem like the only ground truth. The goal posts can keep moving... they can do language but can't implement robots with proper voice or facial expressions, etc. But in theory if there were no more goal posts left, I feel like the architecture argument would ring hollow.
Same with a painting. If an old master draws a wireframe of a dog, people would bid it up at auction and wonder what he meant. If your kid or AI do, no money might change hands. Same output, different context.
So you can’t just use the output, surely?
Can your TI-83 do any proofs?