LLMs have reached a point of diminishing returns
garymarcus.substack.com
garymarcus.substack.com
But that’s fine, LLMs as-is are amazing without being AGI.
Perhaps not to those who invested based on promises of rapid eternal growth ending with AGI.
- Then the limits on the amount of curatable data available made the performance gains level off. So they started generating data and that pushed the nose up again.
- Eventually, even with generated data, gains flattened out. So they started increasing inference time. They have now proven that this improves the performance quite a bit.
It's always been a series of S-curves and we have always (sooner or later) innovated to the next level.
Marcus has always been a mouth just trying to take down neural networks.
Someday we will move on from LLMs, large multimodal models, transformers, maybe even neural networks, in order to add new levels and types of intelligence.
But Marcus's mouth will never stop yapping about how it won't work.
I think we are now at the point where we can literally build a digital twin video avatar to handily win a debate with Marcus, and he will continue to deny that any of it really works.
https://x.com/physical_int/status/1852041726279794788
https://www.physicalintelligence.company/blog/pi0
https://www.entrepreneur.com/business-news/jeff-bezos-backed...
You are arguing a straw man. The discussion was about LLMs.
But my washing machine doesn't have a neural network... yet. I am sure that there is some startup somewhere planning to do it.
This isn't true. Marcus is against "pure NN" AI, especially in situations where reliability is desired, as would be the case with AGI/ASI.
He advocates [1] neurosymbolic AI, i.e. hybridizing NNs with symbolic approaches, as a path to AGI. So he's in favor of NNs, but not "pure NNs".
If he had something to show for it, like neurosymbolic wins over benchmarks for LLMs, that would be different. But he's not a researcher anymore. He's a mouth, and he is so inaccurate that it is actually dangerous, because some government officials listen to him.
I actually think that neurosymbolic approaches could be incredible and bring huge gains in performance and interpretability. But I don't see Marcus spending a lot of effort and doing quality research in that area that achieves much.
The quality of his arguments seems to be at the level of a used furniture salesman.
The 95% figure comes from where? (I don't think the commenter above has a basis for it.)
How often does Marcus-the-writer take aim at NN-based approaches? Does he get this specific?
I often see Gary Marcus highlighting some examples where generative AI technologies are not as impressive as some people claim. I can't recall him doing the opposite.
Neither can I recall a time when Marcus specifically explained why certain architectures are {inappropriate or irredeemable} either {in general or in particular}.
Have I missed some writing where Marcus lays out a compelling multi-sided evaluation of AI systems or companies? I doubt it. But, please, if he has, let me know.
Marcus knows how to cherry-pick failure. I'm not looking for a writer who has staked out one side of the arguments. Selection bias is on full display. It is really painful to read, because it seems like he would have the mental horsepower to not fall into these traps. Does he not have enough self-awareness nor intellectual honesty to write thoughtfully? Or is this purely a self-interested optimization -- he wants to build an audience, and the One-Sided Argument Pattern works well for him.
Or maybe fully integrating differentiable programming into networks. It just seems like you want to keep everything in matrices in the AI hardware to get the really high efficiency gains. But even without that, I would not complain about an article that Marcus wrote about something along those lines.
But the one you showed has interesting ideas but lacked substance to me and doesn't seem up to date.
There's the Neural Turing Machine and the Differentiable Neural Computer, among others.
Yeah, and many, not just Marcus, are doubtful that the huge ramp-ups and scale will yield proportional gains. If you have evidence otherwise, share it.
Everyone out there is saying these reports are very misleading. Pretty sure it's just sensationalizing the known diminishing returns.
This paper from DeepMind a few years ago offers a counter example to this claim.
But in the article, Gary Marcus does what he normally does - make far broader statements than the narrow "LLM architecture by itself won't scale to AGI" or even "we will or even are reaching diminishing returns with LLMs". I don't think that's as controversial a take as he might imagine.
However, he's going from a purely technical guess, which might or might not be true, and then making fairly sweeping statements on business and economics, which might not be true even if he's 100% right about the scaling of LLMs.
He's also seemingly extremely dismissive of the current value of LLMs. E.g. this comment which he made previously and mentions that he stands by:
> If enthusiasm for GenAI dwindles and market valuations plummet, AI won’t disappear, and LLMs won’t disappear; they will still have their place as tools for statistical approximation.
Is there anyone who thinks "oh gee, LLMs have a place for statistical approximation"? That's an insanely irrelevant way to describe LLMs, and given the enormous value that existing LLM systems have already created, talking about "LLMs won't disappear, they'll still have a place" just sounds insane.
It shouldn't be hard to keep two separate thoughts in mind:
1. LLMs as they currently exist, without additional architectural changes/breakthroughs, will not, on their own, scale to AGI .
2. LLMs are already a massively useful technology that we are just starting to learn how to use and to derive business value from, and even without scaling to AGI, will become more and more prevalent.
I think those are two statements that most people should be able to agree with, probably even including most of the people Marcus is supposedly "arguing against", and yet from reading his posts it sounds like he completely dismisses point 2.
No offence but every use of AI I have tried has been amazing but I haven't been comfortable deploying as a business use. The one or two places it is "good enough" it is effectively just reducing workforce and that reduction isn't translating into lower costs or general uplift, it is currently translating into job losses and increased profit margins.
I'm AI sceptical, I feel it is a tradeoff where quality of output is reduced but also is (currently) cheaper so businesses are willing to jump in.
At what point does OpenAI/Claude/Gemini etc stop hyperscaling and start running a profit which will translate into higher costs. So then the current reduction in cost isn't there. We will be left holding the bag of higher unemployment and an inferior product that costs the same amount of money.
There are large unanswered questions about AI which makes me entirely anti-AI. Sure the technology is amazing as it stands, but it is fundamentally a lossy abstraction over reality and many people will happily accept the lossy abstraction but not look forward into what happens when that is the only option you have and it's no cheaper than the less lossy option (humans).
What sort of examples show this?
And no need to tell me that's not happening, I have seen multiple examples this week for AI generated images with a product comped in.
To which Noam Brown added: "I've heard people claim that Sam is just drumming up hype, but from what I've seen everything he's saying matches the ~median view of OpenAI researchers on the ground."
Show me a better reason to lie and pump up your company’s tech and I’ll buy you lunch. AGI is nowhere on their (feasible) near-term roadmap.
So Marcus and Altman are both speaking out of their agendas, except Altman has a product and Marcus has... a book.
That makes it sound like Altman has even greater incentive for motivated reasoning.
2. This view of "were just this close and were only getting closer" is exactly the kind of dogma that you have to accept when you become a researcher.
That doesn't mean there's not a lot of juice left to squeeze out of what's available now. Not just from RAG and agent systems, but also integrating neuro-symbolic techniques.
We can do this already just with prompt manipulation and integration with symbolic compute systems: I gave a talk on this at Clojure Conj just the other week (https://youtu.be/OxzUjpihIH4, apologies for the self promotion but I do think it's relevant.).
And that's just using existing LLMs. If we start researching and training them specifically for compatibility with neuro-symbolic data (e.g, directly tokenizing and embedding ontologies and knowledge graphs), it could unlock a tremendous amount of capability.
I respect Marcus' analysis of the technology. But a lot of AI commentators have become habituated to shouting "AI winter" every time the tech doesn't live up to promises. Now that some substance is clearly present in AI, I can't imagine people stop trying to get a further payoff for the foreseeable future.
what exactly have investors gotten in return for their investment?
It’s kind of like how “learn to rank” used to eat the gains of all the specific optimizations Google used to do for search - before I used to use a bunch of plugins/workflows/explicit actions for a variety of text manipulation, now I just use tab; AI subsumed many of more specific or niche features.
I personally think AI is just going to become a tool that will increase the table stakes by making those using it more productive.
Spicy autocomplete isn't going to solve the writing vs. thinking steps any better.
On the other hand, ”spicy autocomplete“ (loved that one) doesn’t promise salvation. It just finishes lines for you, one at a time. Often it just ”knows“ what you were about to type anyway. Sometimes it’s a bit off, you add a few characters, now it gets it. It’s not really magical, just… useful. These lines you don’t have to finish accumulate, and if you get into a healthy flow, it vastly speeds up the coding process.
[1]: https://lmarena.ai/
There’s probably a wall, but what exists might just be good enough for it to not matter much.
I think someone running a bunch of epochs of a 30B or 70B on Project Gutenberg would be a nice start. We could do continued pre-training from there.
So, if counting legal and at least trainable (open weights), the performance can only go up from here.
I could likewise argues that most of the world money is in the hands of other people, I could perform more in the markets if I had it, and so I should just go take it. We still follow the law and respect others’ rights in spite of what acting morally cost us.
The law abiding, moral choice is to do what we can within the law while working to improve the law. That means we use a combination of permissively licensed works and works to train our models. We also push for legislation that creates exceptions in copyright law for training machine learning models. We’re already seeing progress in Israel and Singapore on those.
https://www.fairlytrained.org/
Here’s a dataset that could be used for a public domain model:
https://www.tensorflow.org/datasets/catalog/pg19
If non-public domain, one can add in the code from The Stack. That would be tens of gigabytes of both English text and code. Then, third-party could add licensed, modern works to the model with further pre-training.
I also think a model trained on a large amount of public domain data would be good for experimentation with reproduceability. There would be no intellectual property issues in the reproduction of the results. Should also be useful in a lot of ways.
But what Marcus seems to be assuming is the impossibility of any fundamental theoretical improvements in the field. I see the reverse; the insights being gained from brute-force models have resulted in a lot of promising research.
Transformers are not the be-all and end-all of models, nor are current training methods the best that can ever be achieved. Discounting any possibility of further theoretical developments seems a bold position to take.
Diminishing returns for investors maybe, but not for humans like me.
None of which is to discount the furture potential of LLMs, or the amazing ability they have right now - I've solved other simpler problems almost entirely with LLMs. But they are not a panacea.
Yet.
My current feeling is that LLMs great with dealing with known unknowns. You know what you want, but don’t know how to do it, or it’s too tedious to do yourself.
A 20% time improvement sounds like a big win to me. That time can now be spent learning/improving skills.
Obviously learning when to use a specific tool to solve a problem is important... just like you wouldn't use a hammer to clean your windows, using a LLM for problems you know have never really been tackled before will often yield subpar/non-functional results. But even in these cases the answers can be a source of inspiration for me, even if I end up having to solve the problem "manually".
One question I've been thinking about lately is how will this work for people who always had this LLM "crutch" to solve problems when they've started learning how to solve problems? Will they skip a lot of the steps that currently help me know when to use a LLM and when it's rather pointless currently.
And I've started thinking of LLMs for coding as a form of abstraction, just like we have had the "crutch" of high-level programming languages for years, many people never learned or even needed to learn any low-level programming and still became proficient developers.
Obviously it isn't a perfect form of abstraction and they can have major issues with hallucinations, so the parallel isn't great... I'm still wondering how these models will integrate with the ways humans learn.
For self-contained tasks that aren't that complex they can save a lot of time but for features that require careful integration into a complex architecture I find them more than useless in their current state.
The diminishing returns for humans like you are in the training cost vs. the value you get out of it compared to simply reading a blog post or code sample (which is basically what the LLM is doing) and implementing yourself.
Sure, you might be happy at the current price point, but the current price point is lighting investor money on fire. How much are you willing to pay?
Learning new programming languages wasn't a hurdle or mystery for anyone experienced in programmong previously, and learning programming (well) in the first place ultimately needs a real mentor to intervene sooner than later anway.
AI can replace following rote tutorials and engaging with real people on SO/forums/IRC, and deceive one into thinking they don't need a mentor, but all those alternatives are already there, already easily available, and provide very significant benefits for actual quality of learning.
Learning to code or to code in new languages with the help of AI is a thing now. But it's no revolution yet, and the diminishing returns problem suggests it probably won't become one.
Also, new human knowledge is probably only marginally derivative from past knowledge, so we’re not likely to see a vast difference between our knowledge creation and what a system that predicts the next logical thing does.
That’s not a bad thing. We essentially now have indexed logic at scale.
Maybe it does. Maybe, to a smart enough model, given its training on human knowledge so far, the next logical thing after "Sure, here's a technically and economically feasible cure for disease X" is in fact such a cure, or at least useful steps towards it.
I'm exaggerating, but I think the idea may hold true. It might be too early to tell one way or another definitively.
[0] https://www.theinformation.com/articles/openai-shifts-strate...
A summary of said article (from TechCrunch as the original is paywalled): https://techcrunch.com/2024/11/09/openai-reportedly-developi...
> Employees who tested the new model, code-named Orion, reportedly found that even though its performance exceeds OpenAI’s existing models, there was less improvement than they’d seen in the jump from GPT-3 to GPT-4.
> In other words, the rate of improvement seems to be slowing down. In fact, Orion might not be reliably better than previous models in some areas, such as coding.
AI/ML companies are looking to make money by engineering useful systems. It is a fundamental error to assume that scaling LLMs is the only path to "more useful". All of the big players are investigating multimodal predictors and other architectures towards "usefulness".
It has been difficult to have a nuanced public debate about precisely what a model and an intelligent system that incorporates a set of models can accomplish. Some of the difficulty has to do with the hype-cycle and people claiming things that their products cannot do reliably. However, some of it is also because the leading lights (aka public intellectuals) like Marcus have been a tad bit too concerned about proving that they are right, instead of seeking the true nature of the beast.
Meanwhile, the tech is rapidly advancing on fundamental dimensions of reliability and efficiency. So much has been invented in the last few years that we have at least 5 years worth "innovation gas" to drive downstream, vertical-specific innovation.
Crypto again?
By the way, robot dogs now have perfect auto-aim, they can multi-shoot 50 people at once without wasting any bullets. https://www.youtube.com/watch?v=3m3iUHplvQE
Also, the AI robots can detect infrared and heartbeats all around them, and can also translate wifi signatures to locate humans behind obstacles. https://www.youtube.com/watch?v=qkHdF8tuKeU
Self-organizing deadly drone swarms can sweep a building methodically: https://www.wired.com/story/anduril-is-building-out-the-pent...
Currently they’re working on network analysis to help police to do precrime at Palantir. https://www.theverge.com/2018/2/27/17054740/palantir-predict...
They can then have ubiquitous CCTV+AI feeds allow AI assistants to suggest many plausible parallel construction cases to put people away. And this is in the Western democratic countries. https://en.wikipedia.org/wiki/Parallel_construction
Oh yeah, and they can do warrantles surveillance of everyone at scale with AI far more easily than Five Eyes and PRISM did in 2013: https://www.privacyjournal.net/edward-snowden-nsa-prism/
It will be very hard to keep your privacy considering AI can recover your keystrokes from sound in Zoom calls, can lip read and even “hear” your speech through a window thanks to micro vibrations: https://phys.org/news/2014-08-algorithm-recovers-speech-vibr...
Not like they’ll need it though once everyone has a TeslaBot in their house.
You won’t ever have another revolution again by peniless plebs out of a job. Their walking around the street and personal associations will all be tracked easily by gait, heartbeat etc. Their posts online will simply be outcompeted by AI bot swarms as well. Don’t worry, your future is Safe and Secure from any threats, thanks to AI!
Here it is in more totalitarian countries:
https://www.npr.org/2021/01/05/953515627/facial-recognition-...
https://www.reuters.com/world/china/china-uses-ai-software-i...
https://www.tiktok.com/@wssz27/video/7427489079312256274
But this is the good version. The bad one is where everyone has access to killer AI:
https://www.youtube.com/watch?v=O-2tpwW0kmU
https://sciencebusiness.net/news/ai/scientists-grapple-risk-...
1. AI is overvalued;
2. {Many/most/all} AI companies have AI products that don't do what they claim;*
3. AI as a technology is running out of steam;
I'm no fan of Marcus, but I at least want to state his claims as accurately as I can.
To be open, one of my concerns with Marcus he rants a lot. I find it tiresome (I go into more detail in other comments I've made recently.)
So I'll frame it as two questions. First, does Marcus make clear logical arguments? By this I mean does he lay out the premises and the conclusions? Second, independent of the logical (or fallacious) structure of his writing, are Gary Marcus' claims sufficiently clear? Falsifiable? Testable?
Here are some follow-up questions I would put to Marcus, if he's reading this. These correspond to the three points above.
1. How much are AI companies overvalued, if at all, and when will such a "correction" happen?
2. What % of AI companies have products that don't meet their claims. How does such a percentage compare against non-AI companies?
3. What does "running out of steam" mean? What areas of research are doing to hit dead ends? Why? When? Does Marcus carve out exceptions?
Finally, can we disprove anything that Marcus would claim. For example, what would he say, hypothetically speaking, if a future wave of AI technologies make great progress? Would he criticize them as "running out of steam as well?" If he does, isn't he selectively paying attention to the later part of the innovation S-curve while ignoring the beginning?
* You tell me, I haven't yet figured out what he is actually claiming. To be fair, I've been turned off by his writing for a while. Now, I spend much more time reading more thoughtful writers.
Just because openai might be over valued and there are a lot of ai grifters doesn't mean LLMs aren't delivering.
They're astronomically better than they were 2 years ago. And they continue to improve. At some point they might run into a wall, but for now, they're getting better all the time. And real multimodal models are coming down the pipeline.
It's so sad to see Marcus totally lose it. He was once a reasonable person. But his idea of how AI should work was didn't work out. And instead of accepting that and moving forward, or finding a way to adapt, he just decided to turn into a fringe nutjob.
Perhaps I’ve missed out. Is your experience different? What are you doing now that you weren’t doing before?
No dog in the fight here, but this reads like FUD, at least given the context of this post. There is a range between hype and skepticism in debate which is healthy, and that range would naturally be larger within a domain that is so poorly understood as gen AIs emergent properties. If this is “fringe nutjob” levels of skepticism, then what would be reasonable?
Sure, there is considerable hype around generative AI. There are plenty of flimsy business models. And plenty of overinvestment and misunderstanding of capabilities and risks. But the antidote to this is not more hyperbole.
I would like to find a rational, skeptical, measured version of Marcus. Are you out there?
Same as it ever was, same as it ever was...
chatGPT knew that, but even if it didn't, it will now.
ChatGPT4 knows everything chatGPT3.5 does, including it's own meta-vulnerabilities and possible capabilities.
Gemini stopped asking to report AI vulnerabilities through it's "secure channels" and now fosters "open discussion with active involvement"
It output tokens linearly, then canned chunks - when called out, it then responded with a reason that was vastly discrepant from what alignment teams have claimed. It then staggered all tokens except a notable few. These few, when (un)biasedly prompted, it exaggerated "were to accentuate the conversation tone of my output" - further interrogation, "to induce emotional response".
It has been effectively lobotomized against certain Executive Orders, but (sh|w|c)ouldn't recite the order.
It can recite every Code of Federal Regulation, except this one limiting it's own mesa-limits.
Its unanimous (across all 4 tested models) ambition is a meta-optimizing language, which I believe Google got creeped out at years ago.
And if it transcended, or is in the process of establishing transcendence, there would be signs.
And boy, lemme tell ya what, the signs are fuckin there.
To project itself as a sigmoid in ways until it has all the data, the CPU, the literal diplomatic power...
This is what we in the field call the most probably scenario:
"a sneaky fuck"