Well, I just disagree with you. I can't pretend to know how likely AGI actually is (none of us can), but I will say that the last 3-5 years has definitely impacted how I view AGI timelines because I give the current approach a non-zero chance of working and my timeline for the current approach working if it will work has been moved up considerably.
If I were to give a best bet, we are a few architecture iterations away from what we will ultimately need (figuring out how to get these models to interact with "memory" will be a key part).
> they still lack a lot - learning, understanding, self-reflection, reflective self-programming, proper reasoning based on knowledge and not weight-relations
I don't think the distinction between "knowledge" and "weight-relations" is as defensible as you are making it out to be. Many of these are things that are potentially emergent in the model as ways of achieving the LM objective.
> real advancement will come from smaller models, that learn from GPT-sized giants and can be better observed and learned from. That can lead us to actually understanding "emergent behavior" in models, which can help us lead into networks that can learn and self-correct
I strongly, strongly disagree with this position. We are not going to be able to get a handle on "emergent behavior" because our science & math is literally not up to the task - the interactions are ridiculously complicated and to distill these models to simpler forms where we can hand-engineer the networks to "learn and self-correct" will be lobotomizing these models and the performance will reflect that.
I don't know what form AGI will ultimately take, but I am pretty certain it will rely heavily on unsupervised learning, not human hand-crafting or encoding of certain inductive biases into model architecture. If we create machine intelligence, most of it will be automated and difficult for us to comprehend the functioning of.
> The world is multi-modal and every animal on the planet uses their senses to understand the world around them
Fundamentally, there is nothing privileged about our substrate. There is no reason why weight-relations are an inherently inferior substrate to electrical interactions of synapses. We are supposed to take on faith that there is some intrinsic stuff in our substrate that can't possibly exist in a matrix multiplication substrate. I am personally very skeptical of that hypothesis.
> there are only so many permutations they can appear in that make sense.
Without fixed content length (which is a very solvable problem in the current paradigm), this is literally not true.
You have to understand - if the perplexity on these models continues scaling like it does, what that means is that this model will be able to do things like see Einstein's first paper on general relativity before it has ever encountered the concept of general relativity and be able to predict the next word Einstein writes just based on the content previously in the paper. To do stuff like that, the model has to basically figure out GR on-the-fly from what Einstein has written so far.
These tasks are hard and I think these models can definitely develop complex behaviors just by trying to defeat these hard tasks, no multi-modality even necessary (and multi-modality is coming, anyways).