If you take the materialist view - and in this business you obviously have to - then the question is "do we have human level computational capacity with some level of real-time learning ability?"
So the yardstick we need is (1): what actually is the computational capacity of the human brain? (i.e. the brain obviously doesn't simulate it's own neurons, so what is a neuron doing and how does that map to compute operations?) and then (2) - is any computer system running an AI model plausibly working within those margins or exceeding them?
With the second part being: and can that system implement the sort of dynamic operations which humans obviously must, unaided? i.e. when I learn something new I might go through a lot of repetitions, but I don't just stop being able to talk about a subject for months while ingesting sum content of the internet to rebuild my model.
We don't build airplanes by mimicking birds. AGIs computational capacity won't be directly comparable to human computation capacity that was formed by messy, heuristic evolution.
You're right that there is some core nugget in there that needs to be replicated though, and it will probably be by accident, as with most inventions.
Hardware is still progressing exponentially, and performance improvements from model and algorithmic improvements have been outpacing hardware progress for at least a decade. The idea that AGI will take a century or more is laughable. Now that's borderline deluded.
We don't build airplanes by having them flap wings, but we do build planes which are considerably more aerodynamically efficient then a bird. i.e. we considerably exceed the underlying performance metric governed by the physics.
Which is my point: an AI need not work anything like a human, but it seems obvious that model training taking months of time to incorporate new information fundamentally cannot be close to AGI since even far simpler creatures can learn and retain information on much shorter timescales.
I'd give the notion more credence if we plausibly had say, a model with a context window which could hold a day of information, and then do could train overnight on that such that it would incorporate that context window into the next day of activity.
My prediction: AGI will come from a strange place. An interesting algorithm everyone already knew about that gets applied in a new way. Probably discovered by accident because someone—who has no idea what they're doing—tried to force an LLM to do something stupid in their code and yet somehow, it worked.
What wouldn't surprise me: A novel application of the Archimedes principle or the Brazil nut effect. You might be thinking, "What TF to those have to do with AI‽ LOL!" and you're probably right... Or are you?
What's the fundamental, absolutely insurmountable capability gap between the two? What is it that can't be bridged with architectural tweaks, better scaffolding or better training?
I see a lot of people make this "LLMs ABSOLUTELY CANNOT hit AGI" assumption, and it never seems to be backed by anything at all.
As the most naive implementation: the LLM just collects the relevant training data, and then runs PEFT on itself. Then it tests the tuned version of itself to see whether the tune was any good and is worth retaining.
Sure, that naive approach would require an LLM to be good at assembling training sets, and validating performance gains, and it would be very computationally expensive. But none of those things would somehow make it not-an-LLM-anymore.
A part of it, perhaps: I think of it like 'computer vision'; LLMs offer 'computer speech/language' as it were. But not a 'general intelligence' of motives and reasoning, 'just' an output. At the moment we have that output hooked up to a model that has data from the internet and books etc. in excess of what's required for convincing language, so it itself is what drives the content.
I think the future will be some other system for data and reasoning and 'general intelligence', that then uses a smaller language model for output in a human-understood form.