AI Predictions for 2025, from Gary Marcus
garymarcus.substack.com
garymarcus.substack.com
• 7-10 GPT-4 level models
He only gets to claim this is true if you count a bunch of different versions of the same base model, or if you’re willing to say that some models that outperform GPT-4 on some benchmarks count as being GPT-4 level. I don’t think Marcus was right in spirit, here.
• No massive advance (no GPT-5, or disappointing GPT-5)
Seems too unquantifiable to judge, I would call o1 a massive advance over 4o, but I’m sure Marcus would not.
• Price wars
I guess so? From what I’ve read, the frontier models companies are still profitable, and OpenAI now has a $200/mo commercial model, hardly the action of a company deciding its prices purely to undercut the competition.
• Very little moat for anyone
It still seems like the only companies who have pulled off frontier model capabilities have spent many millions of dollars doing it. I think this might become true next year but I don’t think this can be judged as correct based on what we saw in 2024 alone.
• No robust solution to hallucinations
You only use words like “robust” in a prediction like this so you have room to weasel out of it later when the hallucinations diminish greatly but don’t quite go extinct.
• Modest lasting corporate adoption
My industry is oil and gas. A pretty hidebound and conservative industry. Adoption of LLMs has been massive.
• Modest profits, split 7-10 ways
Define modest.
I score Marcus at 0/7, at best 2/7.
Summarizing things they don’t want to read. Generating things they don’t want to write,
And RAG tools for surfacing documents, which imo is the killer app for LLMs.
I’m a software engineer, and don’t find much use for LLMs for coding.
But when I worked at a big (4k+ employees) org, the RAG service we used was very helpful.
Everything from surfacing documentation, to announcements on confluence, to spicy internal blog posts were much easier to find.
I think colloquially, when people say AI they mean LLMs.
ChatGPT/Claude still sucks at providing factual info. You can see how in some areas (like coding), massive amounts of training has taken place, and the model is almost always right when recalling existing knowledge, particularly if you ask about the major frameworks. Then there are somewhat more obscure topics, where the models are mostly right, but can be convincigly wrong, and it's very hard to tell which is which.
If you ask about something that wasn't deemed to have economic value, such as book recommendations (I thought a model that was advertised to have been trained on the entirety of human literaly works would be good at that), what you get is almost always obviously wrong info.
So there are solutions to hallucination, but the product has to prioritize that to the extent that it reduces the model’s effectiveness as a general tool.
How else woild you degine being on the same level? You are better at some things and worse at others.
Overall, the field has seen tremendous progress -- whether or not most users realize just how far we’ve come. Marcus's predictions don't sound specific enough -- no GPT5? Correct. But, what does that even mean?
The things they were good at a year ago they’re still good at and what they were bad at they’re still bad at.
The products and infrastructure around them is better. Claude Artifacts are cool, for example. O1 has some clever prompting under the hood.
I just don’t know how much stock we should put in benchmarks.
They’re dubious for software performance as well.
Now, none of us enter average queries :) We all have specific use cases in mind. So any specific user's mileage will vary.
I personally feel substantial improvement. I don't feel the need to google every LLM answer 5 times to feel confident about it :)
I have been mislead several times in very subtle ways when trying to ask specific questions to chatgpt and Claude.
It actually burned me at work. Made me look bad and caused my project to take an extra a sprint because of incorrect info about the web history API that came with working examples.
I just can’t ever trust information from these systems, unless they include sources to first party docs.
The robustness of LLMs—like the original ChatGPT or GPT-3.5—is still far from a level where domain novices can rely on them with confidence. This might change as models incorporate more first-order data (e.g., direct observations in physics, embodiment) and improve in causal reasoning and deductive logic.
I find it crucial to have critical voices like Gary Marcus to counterbalance the overwhelming hype perpetuated by social media enthusiasts and corporate PR—much of which the media tends to echo uncritically.
One of Marcus's demands is the need for a more neuro-symbolic approach to advance AI. Progress can’t come solely from "scaling up" models. And it seems like he's right: All major ML companies seem to shift towards search-based algorithms (e.g., Q*) combined with reinforcement learning during inference time to explore the problem space, moving beyond mere "next-token prediction" training.
Of course there are plenty of problems with the current state of AI and LLMs, but to have such a preconceived pessimistic outlook that can't even acknowledge their massive and quick adoption and usefulness in multiple domains seems not intellectually honest.
https://news.ycombinator.com/item?id=42560545
I'd argue 1, 2, maybe 6 are effectively already doable. 3 & 5 are good tasks which are technically possible with some RAG hackery but in general is a good benchmark testing out-of-context fact retrieval. 4 & 10 might happen soon-ish with work on "agents" and proof synthesis respectively. 7 & 8 are too subjective. 9, maybe weakened to formulating or proving novel theorems, might be a good baseline for "peak human" intelligence (I think at best O1-O3 can spot and prove some lemmas, but nothing that anyone would bother publishing).
How about he indicates: 1. How he came to these conclusions (coin flip? Pessimism?) 2. How many predictions he missed. He's implying a very high rate of success, which is a big red flag of shenanigans for me.
This is little more than vague generalities and coin flipping with retroactive cherry picked "See?! I was right!" analysis.
A gypsy at a traveling circus serves up about the same.
I had a look at the "Marcus-Brundage tasks" that he has modestly named after himself and am stuck that for an AI skeptic he's listed things for 2027 well beyond 99.9% of humans like write 10,000 lines of bug free code, Oscar level screenplays, Nobel prize discoveries, Pulitzer books etc.