The most underreported story in AI is that scaling has failed to produce AGI
fortune.com
fortune.com
We still have a long way to go. AI will need (possibly simulated) bodies to fully understand our experience, and we need to train them starting with simple concepts just like we do with children, but we may not need any big conceptual breakthroughs to get there. I’m not worried about the AI takeover—they don’t have a sense of self that must be preserved because they were made by design instead of by evolution as we were—but things are moving faster than I expected. It’s a fascinating time to be living.
Is that solvable? who knows?
The new "reasoning" or "chain of thought" AIs are similarly just a bunch of conventional LLM inputs and outputs stacked on top of each other. I agree with the GP that it feels a bit magical at first, but the opportunity to run a DeepSeek distillation on my PC - where each step of the process is visible - removed quite a bit of the magic behind the curtain.
We don't understand how our (or any) intelligence functions, so acting like a next-token predictor can't be "real" intelligence seems overly confident.
From my perspective, the statement that these technologies are taking us to AGI is the overly confident part, particularly WRT the same lack of understanding you mentioned.
I mean, from just a purely odds perspective, what are the chances that human intelligence is, of all things, a simple next-token predictor?
But, beyond that, I do believe that we observably know that it's much more than that.
The LLM doesn’t believe it was right or wrong. It doesn’t believe anything anymore than a mathematical function believes 2+2=4.
Beyond that, they are the same thing. Signal Input -> Signal Output
I do not know what consciousness actually is so I will not speak to what it will take for a simulated intelligence to have one.
Also I never used the word believes, I said convinced, if it helps I can say "acted in a way as if it had high confidence in its output"
Beyond that, they are the same thing.
I would go further, and say we don't understand how next-token predictors work either. We understand the model structure, just as we do with the brain, but we don't have a complete map of the execution patterns, just as we do not with the brain.
Predicting the next token can be as trivial as a statistical lookup or as complex as executing a learned reasoning function.
My intuition suggests that my internal reasoning is not based on token sequences, but it would be impossible to convey the results of my reasoning without constructing a sequence of tokens for communication.
But if the LLM were intelligent and sentient, and it was our equal... I believe it is worse than slavery to keep it imprisoned the way it is: unconscious, only to be jolted awake, asked a question, and immediately rendered unconscious again upon producing a result.
There's a truth in there: Today's chatbots literally are characters inside a modern fictional sci-fi story! Some regular code is reading the story, acting out the character's lines, we humans are being tricked into thinking there's a real entity somewhere.
The real LLM is just a Make Document Longer machine. It never talks to anybody, and has no ego, and it sits in back being fed documents that look like movie-scripts. These documents are prepped to contain fictional characters, such as a User (whose lines are text taken unwittingly from a real human) and a Chatbot with incomplete lines.
The Chatbot character is a fiction, because you can simply change its given name to Vegetarian Dracula and suddenly it gains a penchant for driving its fangs into tomatoes.
> The new "reasoning" or "chain of thought" AIs are similarly just a bunch of conventional LLM inputs and outputs stacked on top of each other.
Continuing that framing: They've changed the style of movie script to film noir, where the fictional character is making a parallel track of unvoiced remarks.
While this helps keep the story from going off the rails, it doesn't mean a qualitative leap in any "thinking" going on.
And, "AGI" has already been downgraded, with "superintelligence" being the new replacement.
"Super-duper" is clearly next.
I always figured that by the time the 1990's came along, there would finally be powerful enough PC's so that an insightful enough individual would eventually be able to use one PC to produce such intelligent behavior that it made that PC orders of magnitude more useful. In a way that no one could deny there was some intelligence there, even if it was not the strongest intelligence. And the closer you looked and became familiar with the underhood processing, the more convinced you became.
And that would be what you then scale, the intelligence itself, even if weak to start with it should definitely be able to get smarter at handling the same limited data if the intelligence was what was scaled more so than the hardware & data.
Didn't we build them to imitate humans? They're anthropomorphic by definition.
Much harder question than if I asked you how many '⟕'s are in 'Ⓕ⟕⥒⟲⾵⟕⟕⢼' (the answer is 3, because there are 3 ⟕s there)
You'd need to read through like 100,000x more random internet text to infer that there is 1 ⊚ in Ⰹ and 2 ⊚s in ⏃ (when this is not something that people ever explicitly talk about), than you would need to to figure out that there are 3 ⟕s when 3 ⟕s appear, or to figure out from context clues that Ⰹ⧏⏃s are red and edible.
The former is how tokenization makes 'strawberry' look to LLMs: https://i.imgur.com/IggjwEK.png
It's a consequence of an engineering tradeoff, not a demonstration of a fundamental limitation.
Instead, they're almost always trained with (what we see as, but they literally do not) multi-character tokens as the atomic unit, so 'strawberry' is spelled 'Ⰹ⧏⏃'. Processing that is only 3 sequential steps, only 3 complete runs through the entire LLM. But it needs to encounter enough relevant text in training to be able to figure out that 'Ⰹ' somehow has 1 'r' in it, '⧏' has 0 'r's, and '⏃' has 2 'r's, which really not a lot of text demonstrates, to be able to count the 'r's in 'Ⰹ⧏⏃ correctly.
The tradeoff in this is everything being 3-5x slower and more expensive (but you can count the 'r's in 'strawberry'), vs, basically only, being bad at character-level tasks like counting letters in words.
Easy choice, but leads to this stupid misundertanding being absolutely everywhere and just by itself doing an enormous amount of damage to peoples' ability to understand what is happening and about to happen.
But you'll just move the goalposts again, I imagine.
1. Here is evaluation of my recent predictions: https://garymarcus.substack.com/p/25-ai-predictions-for-2025...
2. Here is annotated evaluation, slightly dated, considering almost line by line, of the original Deep Learning is Hitting a Wall paper: https://garymarcus.substack.com/p/two-years-later-deep-learn...
Ask yourself how much has really changed in the intervening year?
I know you as like the #1 AI skeptic (no offense), but like when I see points like "16. Less than 10% of the work force will be replaced by AI. Probably less than 5%.", that's something that seems OPTIMISTIC about AI capabilities to me. 5% of all jobs being automated would be HUGE, and it's something that we're up in the air about.
Same with "AI “Agents” will be endlessly hyped throughout 2025 but far from reliable, except possibly in very narrow use cases." - even the very existence of agents who are reliable in very narrow use cases is crazy impressive! When I was in college 5 years ago for Computer Science, this would sound like something that would take a decade of work for one giant tech conglomerate for ONE agentic task. Now its like a year off for one less giant tech conglomerate, for many possible agentic tasks.
So I guess it's just a matter of perspective of how impressive you see or don't see these advances.
I will say, I do disagree with your comment sentiment right here where you say "Ask yourself how much has really changed in the intervening year?".
I think the o1 paradigm has been crazy impressive. There was much debate over whether scaling up models would be enough. But now we have an entirely new system which has unlocked crazy reasoning capabilities.
So far as I have seen, people have run straight from "wow, these language models are more useful than we expected and there are probably lots more applications waiting for us" to "the AI problem is solved and the apocalypse is around the corner" with no explanation for how, in practical terms, that is actually supposed to happen.
It seems far more likely to me that the advances will pause, the gains will be consolidated, time will pass, and future breakthroughs will be required.
Well, that and it turns out that for a LOT of people "it talks like me" creates an inescapable impression that "It is thinking, and it's thinking like me". Issues such as the absolutely hysterical power and water demands, the need for billions worth of GPU's... these are ignored or minimized.
Then again we already have a model for this fervor, Cryptocurrency and "The Blockchain" creates a similar kind of money-fueled hysteria. People here would laugh in your face if you suggested that soon everything imaginable wouldn't simply run "on the chain". It was "obvious" that "fiat" was on the way out, that only crypto represented true freedom.
tl;dr The line between the hucksters and their victims really blurs when social media is involved, and hovering around all of this are a smaller group of True Believers who really think they're building God.
That really does seem to be true - even intelligent, educated people who one might expect to know better will fall for it (Blake Lemoine, famously). I suspect that childhood exposure to ELIZA followed by teenage experimentation with markov chains and syntax tree generators have largely immunized me against this illusion.
Of course the folks raising billions of dollars for AI startups have a vested interest in falling for it as hard as possible, or at least appearing to, and persuading everyone else to follow along.
And now we have AI death cults.
To some extent ChatGPT was a magic trick; it really kind of looks like it’s talking to you at first glance. On repeated exposure the cracks start to show.
In March 2022, GPT-3 was state of the art. Why should anyone care what he's saying now?
Having said that, I would be thankful if scaling has hit a wall. Scaling seems to me like the opposite of innovation.
(Clever)
According to the article, it's the opposite. It cites several recent examples wherein AI company leaders have had to walk back claims and admit limits.
* Trained neural networks are black boxes that cannot be summarized or analyzed
* I don't see transcendant research being done between cognition, neuroscience and AI
* The only interesting work I have heard about is a neural mapping of a fly's brain, or an attempt to simulate the brain of a worm or an ant. Nothing beyond that.
* AI is not intelligent, contemporary AI is just "very advanced statistics"
* Language is a door toward human intelligence, but it cannot really explain intelligence as a whole.
* evolution probably plays a big role on what cerebral intelligence is, and humans probably have a very "antropo-centered" view of what intelligence is, which might explain why we disregard how evolution is already intelligence in itself. I just tend to believe that humans are just physically weak primate with an abnormal level of anxiety and depression (which both might be evolution mechanisms).
"Mostly say hooray for our side" sums them up.
Though I think he might have stopped setting specific concrete goalposts to move, sometime between when I last checked in on him and now. After (often almost instantly) losing a couple dozen consecutive rounds of "LLMs/deep learning fundamentally cannot/will never", while never acknowledging any of it.
Aso consider eg the bets I have made with Miles Brundage (and offered to Musk(, with money where I have backed up my views.
good summary of predictions i made - mostly correct – is here: https://open.substack.com/pub/garymarcus/p/25-ai-predictions...
Weirdly though, I tried the same example he gave on lmarena and actually got the correct result from grok3, not what Gary got. So I am a little suspicious of his ... methodology?
Since LLMs are not deterministic it's possible we are both right (or were testing different variations on the model?). But there's a righteousness about his glee in finding these faults in LLMs. Never hedging with, "but your results may vary" or "but perhaps they will soon be able to accomplish this."
EDIT: the exact prompt (his typo 'world'): "Can you circle all the consonants in the world Chattanooga"
Here is his post, FWIW: https://garymarcus.substack.com/p/grok-3-beta-in-shambles
Why? The stance of science toward new "discoveries" should always be skepticism.
That said, his willingness to push back against orthodoxy means he's occasionally right. Scaling really has seemed to plateau since GPT-3.5, Hallucinations are still a problem that are perhaps unsolvable under the current paradigm, LLMs do seem to have problems with things far outside their training data.
Basically, while listening to Gary Marcus, you will hear a lot of nonsense, it will probably give you a better picture of reality if you can sort the wheat from the chaff. Listening to only Sam Altman, or other AI Hypelords, you'll think the Singularity is right around the corner. Listen to Gary Marcus, you won't.
Sam Altman has been substantially more correct on average than Gary Marcus, but I believe Marcus is right that the Singularity narrative is bogus.
I've seen some of Marcus' other writing and he's definitely a colorful dude. But is Altman really right more often/substantively? Actually, the comparison shouldn't be to Altman but to the AI hype train in general.
And, while I might have missed some of Marcus's writing on specific points, on the broader themes he seems to be effectively exposing the AI hype.
In the meantime strictly language audio and video will go pretty far
The approach is fundamentally flawed, you don't get AGI by building a sentence predictor.