(This follows the pattern I've noticed where the term "AGI skeptic" will soon, if not already, mean "I don't trust AGIs in positions of authority or power" rather than "I don't think the technology is capable of matching our level of cognition.").
So let's grant him 8/9 of his predictions, and turn our attention to the one that seems to be his 'real' prediction. This specific and direct prediction about GPT is that "There will be no viable robotics applications that harness the serious power of GPTs in any meaningful way." which I mean, maybe he's trying to use the words "viable" or "serious" or "meaningful" to weaken his claim so that it's never wrong. If we assume he's not just making a vacuously weakened prediction, then I wonder if he has seen https://palm-e.github.io/ for example. You could say the prediction is wrong already, or that it will obviously be wrong before 2030, or that his phrasing makes it impossible for the prediction to ever be wrong.
Speaking of when or if his predictions can be shown to be wrong, I thought it was weird that he only says "These predictions cover the time between now and 2030." whereas for his earlier predictions he seems to have made a more formalized way of incorporating his dates into his predictions which he's not using for his GPT predictions for some reason:
---
I specify dates in three different ways:
NIML meaning “Not In My Lifetime, i.e., not until after January 1st, 2050
NET some date, meaning “No Earlier Than” that date.
BY some date, meaning “By” that date.
Sometimes I will give both a NET and a BY for a single prediction, establishing a window in which I believe it will happen.
His prediction #4 rings false to me.
And this popular insistence that GPT has no 'world model' is also false IMHO. If a GPT is trained on real-world sensory data, it will develop a spacio-temporal 'world model', almost by definition.
LLMs are trained on human text. Their 'world model' is the world encoded in that text, rather than as sounds and retina images, etc, and so it's inconsistent, incomplete, and 'unreal' in a way which is easy to expose.
But there is a 'world model' of sorts there.
Or, most likely, it will develop a model of the data it was trained on, and not the world this data came from.
I'm also curious what you mean by "real-world sensory data". So far, deep neural net models have been trained mainly (say 80%) on things like text, images, time series... and that's about it really. What is "real-world sensory data" in that context?
What is your 'world model'? It is, ultimately, associations between sensory inputs, and outputs, and sequences of them. Which, as Descartes pointed out, could all be fake, and maybe not from any real world at all.
I mean the point is what is the difference? Is your brain trained on 'the real world', or on your senses (data) of the real world? Descartes argued that you can't tell if there's a real world at all.
By 'real-world sensor data' I mean that the models can be trained on sensor data from the real world (video, audio, feedback from motor outputs, etc), rather than on text, which is only abstractly related to the real world. I believe this is called the token-grounding problem, and goes back to Searle's Chinese Room thought experiment.
GPT-4 language performance improved after adding vision training.