I take your point. And agree there is still a huge gulf.
> Then the system is likely going to look a lot different than it does now
I think I agree with this too, but to challenge the idea: I would never have thought "a fancy autocomplete architecture" could give rise to something as sophisticated as ChatGPT, its flaws notwithstanding. So I don't feel so confident that further iterations of the idea, or iterations that involve other architectures that are "still obviously fake" won't give rise to results that far more terrifyingly convincing than ChatGPT.
Since we don't really understand these architectures, human or machine, I don't see how that can be used as the criteria. Ever more find-grained versions of "output" seem like the only ground truth. The goal posts can keep moving... they can do language but can't implement robots with proper voice or facial expressions, etc. But in theory if there were no more goal posts left, I feel like the architecture argument would ring hollow.