Scaling will never get us to AGI
garymarcus.substack.com
garymarcus.substack.com
Although, of course he is not alone with his argument. E.g. even Yann LeCun is repeating a similar argument on LLMs, and many other serious researches as well, that we probably need a bit more than just LLMs + our current training method + scaling up. E.g. some model extension which handles long-term memory better (lots of work on this already), or other model architectures (e.g. Yann LeCun proposed JEPA), or online learning (unclear how this should work with LLMs), or also different training criteria, etc. In any case, multi-modality (not just text) is important. Maybe embodiment (robots, interaction with the real world) as well.
(Edit Wow, the votes on this comment go up and down. I guess it's a controversial topic.)
Oh wait. I lost all respect to those hype-men who claimed it's just around the corner.
Before transformers, people claimed that you just need bigger neural network, not something new. It did not work.
Then came transformers, the new thing can again scale up better than previous era. Transformers have obvious limit: the computational complexity of self-attention module scales quadratically with the sequence length. We are trying to get partial fix with sparcity and linear tranformers and other techniques, but it gets harder at every turn.
We make progress, but "just scaling" is not working.
You can come down to SF and take a waymo today. It works great.
Level 5 vehicles do not require human attention―the “dynamic driving task” is eliminated. Level 5 cars won’t even have steering wheels or acceleration/braking pedals. They will be free from geofencing, able to go anywhere and do anything that an experienced human driver can do.
The understanding I have is that scalability of LLMs with data took even their developers by surprise. They kept at it because empirically, so far, it's worked. But nobody assumes it'll keep working indefinitely or that it'll lead to AGI alone. If they did, OpenAI, Google, etc. would have fired most of their researchers and would simply focus everything they have on scaling.
When we all know it isn't, and this can have the side effect of creating a bubble. I would be concerned with this as it can have terrible effects to AI research long-term.
Scaling self-attention has also worked much further than it was initially believed it could work. Initially it was always believed to be the main bottleneck, but this turned out wrong. The feed-forward layers are much more the bottleneck when looking at where most of the compute is spent.
It's still a bit unclear how much further we can take the scaling. We are mostly at the limit of what is financial feasible today but as long as some form of generalized Moore's law continues (e.g. number of transistors per dollar), it becomes financial feasible to scale further. At some point we also hit a limit of available data. But maybe self-play (or variants) might solve this.
I guess most researchers agree that just scaling up further is maybe not optimal (w.r.t. reaching AGI), or also will not be enough, but it's a somewhat open question.
- Better training algorithms for training with lower precision
- fastattention, fastattention 2 for decreasing bandwidth inside the GPU
- ring attention for decreasing bandwidth across GPUs
- algorithmic improvements in handling long context window
- MoE
- Lots of algorithmic improvements in fine tuning / alignment
- GROQ hardware architecture (deterministic hardware, storing all data in SRAM for inferencing instead of using cache hierarchies)
- Improvements in tokenization
So far softmax(K * Q')*V is the only thing that hasn't (yet) been touched.If the question is that ,,just'' improving LLM perplexity by further algorithmic and hardware improvements will lead to AGI, at this point many researchers believe that the answer is yes (and many others that the answer is no :) ).
https://www.linkedin.com/posts/yann-lecun_what-meta-learned-...
One might also imagine that as one of the "godfathers of AI" he feels a bit sidelined by the success of LLMs (especially given above), and wants to project an image of visionary ahead of the pack.
I actually agree with him that if the goal is AGI and full animal intelligence then LLMs are not really the right path (although a very useful validation of the power of prediction). We really need much greater agency (even if only in a virtual world), online learning, innate drives, prediction applied to sensory inputs and motor outputs, etc.
Still, V-JEPA is nothing more than a pre-trained transformer applied to vision (predicting latent visual representations rather than text tokens), so it is just a validation of the power of transformers, rather than being any kind of architectural advance.
The question is rhetorical: I can see you're a PhD student. My advice is to learn to have some respect for the person of others, as you will want them to have respect of your person when, one day, you find yourself saying something that "the community" disagrees with. And if you're not planning to, one day, find yourself in that situation, consider the possibility that you're in the wrong job.
And what is the above article saying that you think should not be taken seriously? Is it not a widely recognised fact that neural nets performance only improves with more data, more compute and more parameters? A five year old could tell you that. Is it controversial that this is a limitation?
_____________________
Maybe I'm also in a bubble, but I was speaking mostly about the people I frequently read from, i.e. lots of people from Google Brain, DeepMind, other people who frequently publish on NeurIPS, ICLR, ICML, etc. Among those people, Gary is usually not taken seriously. At least that was my impression.
But let's not make this so much about Gary: Most of these people disagree with the opinion that Gary shares, i.e. they don't really see such a big need for symbolic AI, or they see much more potential in pure neural approaches (after all, the human brain is fully neural).
Yes, I get it. And the perspective you wanted to give was to not take Marcus seriously because the people you follow on social media say he's not to be taken seriously. That's nothing but a form of collective online bullying that attacks the person and not the opinion, and like I say in my other comment above, shameful.
Consider for a moment the impression that you make when you say that some people you know, when they're not publishing on NeurIPS, are on social media dogpiling on someone who criticises their work. That's not researchers any more, but common social media trolls.
>> But let's not make this so much about Gary: Most of these people disagree with the opinion that Gary shares, i.e. they don't really see such a big need for symbolic AI, or they see much more potential in pure neural approaches (after all, the human brain is fully neural).
To my experience, the majority of neural net researchers don't know anything concrete about symbolic AI, just what they have heard second-hand, usually on social media again, usually by people who disagree with Marcus, who's the most famous proponent of neuro-symbolic AI (NeSy). So whatever opinion they have on NeSy is not an informed opinion.
There's plenty of literature on NeSy which is a bona fide field of research with a conference etc. This year Leslie Valiant was the keynote speaker and Yan LeCun the honoured guest:
https://sites.google.com/view/nesy2023
You really don't have to listen to what Marcus says to form an opinion on NeSy. Btw, I am not with them and I think they're going the wrong way, but at least I know what they're doing. That is much less than can be said about most neural net researchers, who rarely know anything outside their own work besides whatever preprint is trending on X. That's to the detriment of nobody but themselves.
Twitter for the past decade plus has really publicized and amplified these petty, close-minded academic cliques. It's pretty disgusting to watch for someone who was also a PhD student eyeing an academic career once.
I'm not sure how you meant that to be parsed.
1) performance only improves by scaling up those factors, and can't be improved in any other way
OR
2) performance can only (can't help but) get better as you scale up
I'm guessing you meant 1), which is wrong, but just in case you meant 2), that is wrong too. Increased scaling - in the absence of other changes - will continue to improve performance until it doesn't, and no-one knows what the limit is.
As far as 1), nobody thinks that scaling up is the only way to improve the performance of these systems. Architectural advances, such as the one that created them in the first place, is the other way to improve. There have already been many architectural changes since the original transformer of 2017, and I'm sure we'll see more in the models released later this year.
You ask if it's controversial that there is a limit to how much training data is available, or how much compute can be used to train them. For training data, the informed consensus appears to be that this will not be a limiting factor; in the words of Anthropic's CEO Dario Amodei "It'd be nice [from safety perspective] if it [data availability] was [a limit], but it won't be". Synthetic data is all the rage, and these companies can generate as much as they need. There's also masses of human-generated audio/video data that has hardly been touched.
Sure, compute would eventually become a limiting factor if scaling were the only way these models were being improved (which it isn't), but there is still plenty of headroom at the moment. As long as each generation of models make meaningful advances towards AGI, then I expect the money will be found. It'd be very surprising if the technology was advancing rapidly but development curtailed by lack of money - this is ultimately a national security issue and the government could choose to fund it if they had to.
Next year he'll be teaming up with Rudy Giuliani to tout the success of SHRLDU at Four Seasons Landscaping.
The AI community asked GPT-4 to send him an invite, and he accepted.
I feel that we don't understand well enough (scientifically) what "general intelligence" even is, and how it comes to be in humans - to make claims either way.
To me, the only honest answer right now seems to be, that we do not know. We have absolutely no clue how close - or far - we are to AGI.
This same fellow was big on autonomous swarms for problem solving, but when I asked him what the problem autonomous swarms were supposed to solve that you couldn't solve more easily and quickly by a LLM talking to itself, he didn't have an answer.
That’s all it really is
On a somewhat related note, I'm reminded of the game Crysis, which was developed with a custom engine (CryEngine 2) which was famously too demanding to run at high settings on then-existing hardware (in 2007). They bet on the likes of Nvidia to continue rapidly improving the tech, and they were absolutely right, as it was a massive success.
[0] http://www.incompleteideas.net/IncIdeas/BitterLesson.html
It’s also silly.
The crytek case is similar, but there’s a big difference. Crytek was betting that performance of hardware will increase. A lot of AI startups are betting that future LLMs will have an new different, fairly vague, capabilities (AGI, whatever that actually means)
With LLMs there's not not only no clear path to the goal, but there's every reason to think that such a path may not exist. In literally every domain neural networks have been utilized in you reach asymptotic level diminishing returns. Truly full self driving vehicles are just the latest example. They're just as far away now as they were years ago. If anything they now seem much further away because years ago many of us expected the exponential progress to continue, meaning that full self driving was just around the corner. We now have the wisdom to understand that, at the minimum, that's one rather elusive corner.
Is that true, though? I think of "grokking", where long training runs result in huge shifts in generalization, only with orders of magnitude more training after training error seemed to be asymptotically low.
This'd suggest both that there's not that asymptotic limit you refer to - something very important is happening much later - and that there are potentially some important paths to generalization on lower amounts of training data that we haven't yet figured out.
A lower error, after a certain point, does not suggest better responses
Targeted AI applications and virtual "executive assistant" agents are going to be huge though.
It hard to build something that is yet to be even logically defined.
Hoping that AGI will somehow just "emerge" from an inert box of silicone switches (aka a "computer" as we currently know it) is the stuff of movie fantasy ... or perhaps religion.
IMHO, something like "movie fantasy" constitutes the core belief of a great many software engineers and other so-called rational people in the tech space.
The ability to generalize completely outside training data isn't that common among humans IME. That is a high bar. How many of us have done so without at least drawing an analogy to some other experience we have had? Truly unique thinkers aren't that commonplace. I myself could have probably just asked an LLM what I would say in this post and gotten pretty close...
In other words, they're in a lab somewhere undergoing perpetual test runs so they don't kill people, so not generally available.
It is like saying that email was available in the '80s, you just couldn't message anybody outside of the university network.
https://www.axios.com/2023/08/29/cities-testing-self-driving...
Waymo remains available to my knowledge.
https://www.theverge.com/23948708/cruise-robotaxi-suspension...
It's a little-known fact that AI means "Actually, Indians."
https://boingboing.net/2024/04/03/amazons-ai-powered-just-wa...
Isn't the second half of that saying that scaling up the dataset was what got them where they are?
When I hear about "scaling up" models, I think there are two parts of it: 1) use the same architecture, but bigger 2) make the training data bigger to make use of all those new parameters.
So when I hear about something that's not scaling, it would require some sort of fundamental change to architecture or algorithms.
In the sense that if you were allowed to run over anyone without legal consequences lots of tech enthusiasts would sign up for this experience?
The obvious flaw in this line of thinking (as related to AGI) is that we/humans can create/engineer the same or similar without any real understanding of the underlying mechanisms involved.
So I don't see how it's a movie fantasy any more than bottling up stars is (well, ICBMs do deliver bottled sunshine... uh, the analogy is going too far here). Anyway, the point is that while brains and intelligence are complicated systems there isn't anything at all known that says it's fundamentally impossible to replicate their functionality or something in the general category. And scaling will be a necessary but perhaps not sufficient component of that, just because they're going to be complex systems.
Maybe ... someday.
But right now, we don't even fully understand the physics that makes a human brain work. Every time someone starts investigating, they uncover surprising new complexity.
https://theness.com/neurologicablog/is-the-brain-analog-or-d...
Yes, it only took a few billion years of trial and error. And it didn't manage to do it with inanimate objects either.
We know that the human brain is organic, mostly analog and much more complex than a digital computer.
If analog processing were relevant (it most likely isn't) then Artifical Neural Networks could also be implemented on analog silicon circuits.
As for the complexity, well, the systems are getting more complex. Straight lines on log charts. That's what scaling is about. Getting there, to that complexity.
So I'm just not seeing any knockout argument from you. We can't predict the future with certainty. But the factors you present do not appear to be fundamental blockers on the possibility. You're pointing at the lack of an existence proof... which is always the case before a protoype.
It sounds like you're basically abandoning forward-thinking and will only acknowledge AGI when it hits you over the head. A sign of intelligence is also the ability to plan for an unseen future.
So evolution came up with solutions for surviving on planet Earth, which isn't necessarily the same as general problem solving, even though there are significant overlaps. Just the $.02 of a layperson.
The equivalent of a human brain just "emerging" from inert silicon switches without any real understanding of how or why --- that's PFM (Pure Friggin' Magic). There is no logical reason to believe it is even possible or practical --- yet people still believe.
It's the modern day version of alchemy --- trying to create a fantastical result from a chemical reaction before it was understood that nuclear physics is the mechanism required. And even with this understanding, we have yet to succeed at turning lead into gold in any practical way.
Well, sure. But then, we never observed nature turning lead to gold in any cheap way, whereas we do observe nature running complex intelligence on relatively cheap hardware (our brains).
There's a pretty big difference between trying to do something never Nature did (alchemy), versus trying to replicate in a kinda-parallel kind of way something we've observed Nature doing (intelligence).
You might even agree that not a single directed thought went into the design of that whole system.
Having established that, it seems laughable to me that it would be an "impossible fantasy" to replicate this with purpose-built silicon.
Sure, our current whole-system understanding could still be drastically improved, and the sheer scale of the reference implementation (human connectome) is still intimidating even compared to our most advanced chips, but I have absolutely zero doubt that AGI is just a matter of time (and continuous gradual progress).
If a lack of understanding did not stop nature, why should it stop us? :P
To me, all the arguments against AGI appear unconvincing, motivated by religion/faith, or based on definitional sophistry.
But I'm very open to have that view changed...
Religion is belief in the unknown without reason or logic.
Do you know of any inanimate object that is truly "intelligent"?
Without ever knowing or seeing a single, real working example, you "believe" and have "faith" that it is possible. By so doing, you are practicing religion --- not science.
There is a vast difference between "engineering challenges have not yet been proven moot by a working prototype" and "precluded by physics or mathematics".
The (in)animate distinction is more akin to net-positive fusion in a man-made vessel vs. in a gravity well rather than obeys the 2nd law of thermodynamics vs. perpetuum mobile.
Nope.
Before the Wright brothers, we knew it was scientifically possible to suspend objects heaver than air in an air current. For example, kites and balloons.
The only examples we have of "intelligence" are organic in nature.
Since we have no examples to the contrary; for all we know, "life" and "intelligence" could be somehow inter-related. And we don't currently fully understand or know how to engineer either one.
Yet people have "faith" --- just like the alchemists. "Believe" what you want --- but it's not science.
High-abstraction category-words such as "intelligence" are not fundamental properties of nature. They're made by man. Which means we currently happen to define intelligence in a way that we primarily observe in organic objects. So there's some circular reasoning in your argument. If we look at all the more fundamental building blocks such as information-processing, memory, adaptive algorithms, manipulating environments then we know all of them already are implementable in non-organic systems.
It is not a blind belief decoupled from reality. It is an argument based on the observation that the building-blocks we know about are there and they have not been put together yet due to complexity. There also is the observation that the "putting together" process has been following a trajectory that results in more and more capabilities (chess, go, partial information games, simulated environments, vision, language, art, programming, ...), i.e. there's extrapolation based on things that are already observable. Unless you can point at some lower-level piece that is only available in organic systems. Or why only organic systems should be capable of composing the pieces into a larger whole. I am not aware of any such limiting factor.
Just like "suspending" objects in air was possible and self-propelled machines were possible, even if not self-propelled flight had not yet been done at that time.
To me the null hypothesis is that non-organic intelligence is possible --at worst (!!)-- simply by emulating exactly what an intelligent organic entity does (but it seems obvious to me that this approach is likely to be wasteful and inefficient).
It may not be necessary --- but at this point in time, we really don't know enough to state this as a fact.
What we do know is that our only working examples of "intelligence" are all organic and mostly analog --- not digital. Suggesting that real "intelligence" is possible without "life" --- is unsupported by any evidence and is a pure leap of faith at this time.
> Suggesting that real "intelligence" is possible without "life" --- is unsupported by any evidence and is a pure leap of faith at this time.
Strongly disagree on this; If you accept that brains effect intelligence by selfinteraction accoring to physical laws, then "computability" directly and inevitably follows. And that implies artificial brains are feasible...
Given enough time, money, effort and energy, anything is *feasible* --- even converting lead into gold.
Expecting it to just *emerge* somehow on it's own without any real design or plan or understanding of the scope involved --- that goes way beyond *feasible* or *practical* and heads straight for *magical thinking*.
This seems strawmanny. There has been an absolute ton of research into figuring out what kind of design/scope is necessary. More planning and research is happening all the time, and architecture is specifically being changed with a mind towards specific new functionalities. It's not haphazard; it's planned. Yes, there is some room for 'emergence' or learning, but that doesn't mean there isn't also a lot of structure and planning.
The picture you present of NN research is so far off from what I see actually being done in this field that I'm a little boggled. Are you sure you're in touch with what's actually happening in the field?
So, your argument is that by making something intelligent we would necessarily also make it animate?
That's reasonable, but doesn't impact the plausibility of AGI. “Artificial” is not “inanimate”.
Correction. I'm atheist. I don't see AGI for the future. Same reasons as the poster who responded to you below: evolution may not have had a design, but we know it works well on animated creatures.
We have yet to see any examples of inanimate objects exhibiting any sign of intelligence, or even instinct.
What we do know is that, even after billions of years of evolution, not a single rock has evolved to exhibit a sign of intelligence.
You're saying that with an intelligent hand directing the process, we can do better. I understand the argument, and I concede that it is a reasonable argument to make, I'm just unconvinced by it.
Can you give me a clear, minimal definition for both "intelligence" and "instinct"?
Because to me, instinct seems to be already achieved by existing inanimate control systems for decades now, and "decent understanding of natural human language" (as achieved by todays LLMs) is more than enough to count as intelligent for me.
If we just started simulating neurons in a brain exactly, what would prevent us from achieving inorganic intelligence in your view?
Most people (say, 999999 out of every 1000000) will consider a rock unintelligent and Stephen hawking to be intelligent. You can't call them wrong because "they cannot define intelligence".
> If we just started simulating neurons in a brain exactly, what would prevent us from achieving inorganic intelligence in your view?
Nothing. But the word "If" is doing all the heavy lifting in that argument. I mean, we aren't doing that at the moment, are we? It's not clear that we might ever discover exactly how the collection of neurons we call a brain "exactly works".
IOW, if the human brain was simpler, we'd be too simple to understand it: this may already be the case!
And the evidence for "we may not figure out how the brain works exactly enough to clone the mechanism it uses" is a lot larger than "we might figure out how the brain works well enough to clone the mechanism it uses."
This is why I remain unconvinced. Even though I think your position is a reasonable one to hold, the opposite position is, IMHO, just as reasonable. It's got nothing to do with religion or superstition.
Is it even appropriate to apply "after billions of years of evolution" to rocks? Rocks don't evolve. Evolution can't act on them.
And, we aren't generally looking to "evolve" AI, so it's kinda a moot point there, too.
What is it about animation that, to you, is important for intelligence? Presumably it's the training data + agency, but.. are these not possible without physical-world "animation"?
But also: how come AI can't be animated?
AKA, we know it’s possible for atomic systems to be as intelligent as we are, because we are. We suspect the limit is far beyond us because computers have dramatically faster processing, essentially perfect memory systems, etc
I think AI can easily have more strategy and patience and coordinate 10000 swarm members than you. Why will we need you to do anything at all?
The "scale is all you need" argument, from anyone intelligent, assumes that other obvious deficiencies such as lack of short-term memory (to enable more capable planning/reasoning) will also be taken care of along the way, and I'd not be surprised to see that particular one addressed in upcoming next-gen models.
There's a recent paper here from Google suggesting one way to do it, although may other ways too.
https://arxiv.org/pdf/2404.07143.pdf
The more interesting question is what other components/capabilities needs to be added to pre-trained transformers to get to AGI, or are some of the missing pieces so fundamental that they require a new approach, and can't just be retrofitted along the way as we continue to scale up?
But scaling is necessary if we ever want to get to AGI.
One thing his argument has a point is that we're really close to hitting a ceiling in improvements for LLMs, unless a new groundbreaking research arrives.
Just shoving more data, increasing parameters etc won't get us anywhere. Not to mention LLMs behave more like a data compression algorithm than anyform of "artificial" intelligence.
What I personally find troubling is what will happen to NVidia and other pick sellers once the LLM hype fades away and people are aware of its limitations, will we have still big investments in scaling? Likely not.
We just need to methodically sample the under sampled portions of the source distribution. Not more data, better data.
The problem is that they don't exist in discrete missed data sets they only exist in larger data sets that also contain large amounts of well sampled data.
If you could separate out the existing data you could cherry-pick the under sampled data, but do that requires the very classifier you're trying to build in the first place.
This reaction is like saying "if we knew everything we'd be able to answer every question"
You don’t need a fancy model trained on all languages to tell you its training set underrepresented Finnish. You just need the training set curators to say “hey there’s no Finnish in the inputs”.
If the goal is high quality models, why be a purist about only using the resulting model to measure deficiencies in the training set?
Could you elaborate on this?
If there's one thing humanity doesn't need, it's the kind of person who leads a Silicon Valley corporate entity creating a self-aware consciousness.
But humans also have this limitation. Every science discovery is just some new data we can train on and create models for, both mentally and mathematically.
It would an powerful tool, changing the world but the post would be the same 'nothing to see here folks' post as this.
It will change the world for the worse. We don't need more progress and more technology. What's the point of living if we're going to give all the funnest work to AI? Idiocy.
That makes the posts devoid of meaning. It's gist is unaffected by the things happening in the real world. Will always be "I told you this wouldn't work".