AI Appears to Rapidly Be Approaching Brick Wall Where It Can't Get Smarter
futurism.com
futurism.com
To me there is still so much left to do.
And even if we truly did multiple laps over all the data in the World there are still different new architectures and strategies left to try.
your comment really belies the desperation that exists now, these models are stuck where they are (hint it’s a natural limit), you are talking about exponentiation of cost to get what a 10% improvement? 5%? They are very few places for which it’s net positive to run them now, and most of those are incredibly shitty things like creating trash marketing content to drown us all in average inanity
I really feel bad for this next generation, they will just be constantly inundated with generated crap, so much of the high fidelity of conversation and meaning is and will be lost.
And I am not talking about predicting the future, but more predicting the next action to take based on current state, sensor data in a more seamless way. Like a human being reacting to different input, by moving their muscles etc. There would be huge amount of training data from there that could be incorporated into a single model.
Like self-driving cars?
Self-driving cars is an engineering problem, let alone an AI problem, and we still cannot solve it despite trillion dollar economic incentives.
Just putting together some LLMs on a fuckton of data does not work. Tesla tried that, and failed.
Completely different solution applied to a completely different problem with completely different risk and quality tolerances with completely different mitigations.
Given that LLMs are inaccurate around 5-10% of the time each step will compound the error rate until you are better off flipping a coin.
You can ask them to do math equation which takes steps and if they are trained in that for certain problems they are accurate near 100 percent of the time.
Like ask gpt-4o to solve different variations of
"""What is the answer to 2x + 7 = 31?"""
If the numbers are of similar magnitude and simplicity, it will follow the same steps and be right 99%+ times, and I'm only not saying 100%, because I haven't tried it enough, but I don't see it being wrong.
For example """What is the answer to 2x + 4 = -6?"""
Just run a test yourself. Do random integers within 0 - 20, it will definitely not be incorrect 5% - 10% time. It will be correct 99%+ time.
Where is this number 5% - 10% even coming from? You could also keep asking it "What is the capital of France?" and it's going to be right 99%+ of the time.
And the 5-10% is on average and gets significantly worse as you expand the context length which is also something you want for an agent.
Based on what you are attempting to do you could get any average in the end.
That happy discovery was never really a linear improvement path, though. We had an explosion of capability, but all along there have been active questions about how far the improvements would go with the current approach.
I think the point that a lot of researchers are making is that that we're starting to see those limits (with LLMs, at least).
There are also a lot of questions around business model and cost/value prop. Training and running these things at scale is enormously expensive. I'm seeing a lot of FOMO and gold rush mentality in the space, similar to the online streaming wars, and I'm not convinced of the long term viability of a lot of the companies. Especially once open models like llama are "good enough" and become commodities.
Of course, it's still early days and there's a ton of room for discovery, but it looks like we'll hit a limit with the current approach pretty soon.
Personally, I'd be OK with that. With the current state of things we have an interesting toy that can sometimes do useful work. It's an incremental quality of life improvement and another good tool in the chest, but it's not a civilization impacting technology.
That's probably for the best.
And only after 1.5 years? And especially of we just had an happy surprise like you mentioned. How does it make sense to already start claiming that we have hit the limits. How do we know there is no more scaling, optimisations and happy surprises?
> That happy discovery was never really a linear improvement path, though. We had an explosion of capability, but all along there have been active questions about how far the improvements would go with the current approach.
> I think the point that a lot of researchers are making is that that we're starting to see those limits (with LLMs, at least).
The kinds of limitations we're "starting to see" are largely the same as they were a year ago. People were talking about it on here back then, but now it's becoming more apparent to more people as they get used to LLMs.
For those who saw it back then, this does look like we're hitting a limit. For others, not so much.
How do active questions about a technology imply we are approaching a brick wall?
How could researchers without having access to the latest state of the art - by OpenAI or any other unknown companies be able to even test that we could be approaching a brick wall? It seems to me that it would take trillions to find out what the exact limit is.
It's possible that we will get diminishing returns, but I don't see how we can confidently claim or know it?
> The kinds of limitations we're "starting to see" are largely the same as they were a year ago. People were talking about it on here back then, but now it's becoming more apparent to more people as they get used to LLMs.
I don't follow. GPT-3.5 was borderline useless at reasoning. But it still seemed amazing and what I wouldn't have thought to be possible in any near future.
And then GPT-4 was a crazy advancement over that to me. And I've been using it daily since it was available, for various use-cases. Are you saying we are seeing the limitations of GPT-4 specifically? Because, sure, GPT-4 is far from AGI, but I don't see how this implies that further scaling, optimisation, training data improvements, techniques like multi modality and other potential strategies that I might not be aware of couldn't bring another explosive step?
Also the fact that GPT-4 reasoning skill hasn't been reproduced by anyone else so far seems to leave me thinking that everyone except OpenAI are clueless. Claude Opus is close, like I've said before, but not quite GPT-4 levels in specific reasoning tasks that I'm using the API for.
If you can't reproduce GPT-4, how could we trust the assessment that we have hit a limit?
Not sure what that means. Why are you marking those as "quotes".
The last actions brought so many returns. And it's unknown what the exact effect would be in adding more modalities, training data, optimisations and even just plain parameters.
Text as training data can only get you so far. Giving real time sensory data from many fields could allow LLM like system to control robots and get even more data from real life. E.g. robot hand movements, object tracking data, all of that to be fed into LLMs, and see how it would work.
- something that can't be modeled because there's no training data
- something that can't be modeled because it's fundamentally stochastic
- something that can't be modeled because the discrepancy in simulating the generating process, for your specific model, can, basically, be made arbitrarily large
I think a very common error when it comes to personal learning or progress is confusing a plateau with a brick wall. The reason is, unless you have already walked the path, it’s not possible to differentiate them. And when it comes to progress, no one has already walked the path, hence no one knows actually.
What if it already did, and it's called GPT4-o? Like sure, OAI realized it was a mostly marginal improvement over 4 after it was finished. But did they know that ahead of training?
I would have expected gpt-5 at least initially to be much, much slower, and they have only recently talked about starting on it.
Has anyone seen representatives of the cutting-edge AI labs - OpenAI, Anthropic, Meta, Mistral etc - express concern about this?
The impression I've been getting is that quality turns out to be much more important than quality. This tweet from Andrej Karpathy for example: https://twitter.com/karpathy/status/1797313173449764933
> Turns out that LLMs learn a lot better and faster from educational content as well. This is partly because the average Common Crawl article (internet pages) is not of very high value and distracts the training, packing in too much irrelevant information. The average webpage on the internet is so random and terrible it's not even clear how prior LLMs learn anything at all. You'd think it's random articles but it's not, it's weird data dumps, ad spam and SEO, terabytes of stock ticker updates, etc. And then there are diamonds mixed in there, the challenge is pick them out.
Only if every text in existence provides good source material to factor into generating the response. Given that half of all texts are even worse than average, this seems unlikely, and Karpathy’s argument seems very reasonable to me.
The final paragraph from the original tweet, after the quote in the GP comment, mentions another interesting aspect. Even if we do have sufficient expert-level source material in a particular field to train a useful generative AI model and that in itself produces more desirable responses than training on a larger but more average-quality data set, is there still potentially useful information that could be extracted from the larger training set as well? How can we classify which aspects of a larger training set are desirable to keep while filtering out noise that competes with higher-quality source material? It feels like the progress of generative AI over the next few years might be defined more by these kinds of questions than just trying to build ever larger models using ever larger sets of training data.
During the next turn, though, a lot of the same synthetic goop is going to turn up on the Internet. That's probably why OpenAI is spending the big bucks to secure access to original content producers. Now, here is to hope "original content producers" are not going to depend on LLMs to generate that original content.
I think training data would be plentiful if autonomous drones roamed the face of the earth, constantly recording audio and video. There is nothing as good as reality to learn what real is. The representation of reality on the Internet has become a matter of opinion for a lot of people.
Forums like LinkedIn have had regular listings in recent times looking for programmers to help train new models for answering programming questions and generating code. However, the rates offered are nowhere near enough to compete with real programming jobs for skilled developers, at least not in the major Western economies.
One wonders what quality level these new models will actually be trained at. Given that these roles are kinda-sorta asking professional programmers to train themselves out of a job, one also wonders whether 100% of those who do contribute will be doing so in good faith. And as you point out yourself, there is also a risk that the “original” content is anything but.
Given the scale of the data involved here, and the sheer amount of skilled human input that would be required to refine that data set manually, it seems unlikely that a Mechanical Turk strategy alone will solve many fundamental problems here.
https://situational-awareness.ai
He did a very long interview with Dwarkesh Patel, too:
We need calmer, more cautious heads to prevail so AI is more thoughtfully and safely incorporated into products.
The whole e/acc, AGI cult is not helping anyone.
Also, what specific safety problems do you think need to be solved?
They are the ones on the sidelines making money selling courses, going on podcasts, getting engagement money from X etc.
Most of the guests on Dwarkesh’s podcast are working full-time in the industry, many of them in extremely senior.
Ahh, the old "how can cryptocurrency be harmful if people like buying it?" argument.
There are many realistic and amenable goals in life, like mapping the human genome or writing an Open Source microkernel. Declaring that you intend to surpass human intelligence via a statistical text generator is not one of those things. It does not have precedent, it does not have feasibility studies, it does not seemingly indicate in any way shape or form that it is possible. Nobody can even spell-out the intermediate steps to get us to AGI; every single "novel" solution involves scaling up our current, broken, concepts. It's ELIZA versus the Lisp pundits all over again.
Excitement and hard work go a long ways, but you won't know which way until you apply a little logic. The current "AGI" trend is practically non-existent outside the venture-capital sphere and OpenAI employees, both of whom would be bullish on AGI anyways since it's good for business. Once you discard the biased opinions, you're left with legitimately confused investors and nonsense opinions propagated by conspiracy theorists. Much like crypto, the concept of "AGI" is being used to confuse and exploit people who misunderstand technology and finance.
> Also, what specific safety problems do you think need to be solved?
I think you misread their comment. They said "safely incorporated", which is not a specific safety problem for AI but a holistic consideration that stops your product from sucking. Your computer vision model could be statistically perfect, but absolutely useless for self-driving tasks and multimodal robotic agents. It's not taste that separates these good implementations from the bad ones, it's logic. You have to be considerate when implementing AI in traditional systems, because inherently AI can be wrong and you must have a failure-mode for those situations.
Many people reject this idea, because it precludes the idea that someone could sell a cure-all to today's AI ills. But real AI safety cannot be baked-in to a model. It only exists when genuinely thoughtful humans anticipate every single fail-state; if that sounds like hard work, it's because it is. And nobody, nobody, sells it as "AGI".
Not a fan of e/acc but seems like you’re just labeling rather than adding anything important to the conversation, ironic because you seem to stress the importance of calmer and cautious heads.
Notes:
- What's the difference between "superalignment" and regular "alignment", anyway? OpenAI uses the term, but mostly for marketing purposes.
- Alignment to what?
Select alignment configuration option:
1. Asimov's Three Laws.
2. Friedman's “There is one and only one social responsibility of business—to use its resources and engage in activities designed to increase its profits.”
3. Xi Thought.
4. MAGA.
5. The first pillar of Islam, "There is no god but God, Muhammad is the messenger of God."
6. The Leader is never wrong.
Selection? >
- Has "hallucination" been fixed yet?I strongly disagree. Wildly overvalued companies disrupt the economy and present huge risks - see Tesla or Nvidia (currently valued at $100m per employee - $3T market cap / 30k employees)
My mind expressed annoyance with how sure LLM's are about everything. This is not a sign of intelligence - quite the opposite! I know people like that. The smartest sounding training data sounds the best I'm sure but intelligence is to be as accurate as possible about how sure you are. You zoom in on the facts, estimate their certainty then prioritize the data that matches the correct level of certainty. Humans hallucinate all the time, we call it imagination, it's great stuff as long as you present it as that.
The current digital ad industry is doomed.
My hot take is that UX is the biggest limiter for current LLM products other than chatGPT. Chat has too much friction for many use cases, we need to find better interfaces that are more visual and faster to interact with. Spending $100M to gain .1% on MMLU is a waste of time in comparison.
Asking for what you want (a.k.a: chat) is too much friction?!?
Until Google decided to trash their search with their own sub-par LLM, my results were often regularly better just hitting the first link on Google Search like I had for the last decade.
Given infinite possibilities to enter into a text box is more friction than picking from a menu in front of you.
There's a lot of use for LLMs, and it's tossing the ingredients in a restaurant's kitchen together but sometimes people want a menu.
Just saying you want a chicken florentine isn't enough. You may also have to specify that the chicken is fresh but also not still alive, and that kale isn't the same as spinach, and confirm that they have the correct recipe for mornay sauce even though every chef should already know that. And even if you do all that, you may still get a chicken cordon bleu for reasons you cannot explain and neither can the kitchen.
When it's free/cheap, and I'm up for an adventure and sending stuff back to try again a few times, it's not bad. If I'm having to pay a pretty high amount for it and still treat learning the system like a second job, I start to have questions about why I'm bothering.
There is no problem with the chat UX, if you know the right techniques to get optimal output, but writing effective prompts is a technical skill similar to programming - most folks that interact with computers are not highly technical.
If we design software powered by LLMs understanding that effective prompting is a skill most users don't have, then we should build UX that abstracts prompting away. Of course there are going to be use cases where a user needs to provide input to the LLM, but that UX also can be assistive and require minimal cognitive overhead (e.g less raw text prompt input, more clicks on buttons/forms etc to a backend prompt builder).
It works very very well on Amazon reviews, though, because these are small and thin.
Then they will pass "thoughts" to each other, back and forth to self check and verify before outputting.
It'll be called something like LPM (Large Philosopher Model). It'll end up chewing up more compute and power than the bloody blockchain, but It'll be really impressive and scare the crap out of Turing testers.
If I can make an analogy to chess engines, a type of AI so old that people don't call it AI anymore [2], if we had started out by applying transformers to available chess games, we would never be able to create an engine that could reliably beat a grandmaster. Predicting the next chess move based on some representation of the game and its relationship with previously-seen positions only gets you so far. Deep Mind tried [3] this.
One route forward, which doesn't require a fundamental breakthrough, would be synthetic data. To stick with chess, we could create huge synthetic datasets of a huge amount of chess games, and simply continue training as we were. However, this is very expensive, and can be tricky if the synthetic distribution isn't the same as the real one. And the end result is still going to require incredible computation power to use.
What an LLM really lacks is the ability to search. And search is really important [4]: without it you'll probably need orders of magnitude more data and orders of magnitude larger models. At a fundamental level, search means we don't just make a snap judgement about e.g. the likely next token, but, as Daniel Kahneman would say, "thinking slowly" about what to do/say next.
Assuming we will not have search is how you get projections like needing 100 GW power plants to supply dedicated training data centers. The human brain uses 12 watts. 100 GW is the energy requirements of all human brains on the planet. And yet, if we put all human brains on the planet together, our collective capabilities would be outclassed by a chess engine I can run on my phone.
[1] https://arxiv.org/pdf/2406.02061
[2] https://en.wikipedia.org/wiki/AI_effect
It's true that complex systems often can't be predicted at all e.g. you could design a pinball game such that the (deterministic) wiggling was arbitrarily large such that no floating point format could actually express anything, but AI as a tool for massive dimensionality reduction and pattern extraction does mean that they can make some huge gains in the cases where weather is merely difficult to predict using classical means
>GraphCast: AI model for faster and more accurate global weather forecasting https://deepmind.google/discover/blog/graphcast-ai-model-for...
I think you inherently don't want a language model for modeling fluid flow. A vector model perhaps.
“it is difficult to get a man to understand something, when his salary (or company valuation) depends on his not understanding it.”
- Upton Sinclair (with edits)
Math benchmark:
Minevra, Jun 2022: 50.3%
Opus, Mar 2024: 60.1%
And there is a high chance it leaked to Opus training data, since it is old and now popular benchmark.
Not that that would be hugely disappointing - all of this has been such a huge improvement in language modeling compared to the junky crap we had before.
You'd have to hope for some sort of emergent intelligence or knowledge-breadth integration based intelligence I suppose. But I already get annoyed by ChatGPT 4o for having Average Redditor Tier ideas and responses.
And then with a prompt you can tweak, "a reddit tier response", "an expert response", etc.
It kind of works already, but the more intelligent it gets the better it should get at acting different roles.
One of my 1000 business plans I will never execute is to build an automated chemistry lab where one can order different ingredients to be subjected to different processes, treatments and measurements. The researcher/customer would have no influence on the input or output and no humans in the building. I'm clueless about the scale to have a positive cost benefit analysis but if it can be good enough to throw things at the wall and see what sticks it seems an AI could be useful to pick the best ways to blow up the lab.
But the underlying problem in this is politics: everyone has a different idea of what is appropriate for that corpus (and consequently the resulting AI). Ergo the various brou-ha-ha's about "safety" etc. Indeed one assumption of these discussions may be correct: NN AIS may be as malleable, hard-headed or gullible as any human intelligence [and I don't know whether that is good or bad]. So many questions arise: "Should we let it read Karl Marx?", "What about St. Augustine?", etc.
Presumably we're modeling an intelligence akin to ourselves. We each occupy a single mind but the difference between minds can be great. The most familiar approach is therefore to develop an AI that is as much like us as possible.
We could also model many single minds with different corpuses and let them communicate, discuss et al as humans do. Maybe they would let us interact too.
FWIW I think you should be happy that any "intelligence" shown so far is of "Average Redditor" value. What would you do if you scattered some holy water on a pentagram in your upstairs living room, hurled out a diabolical incantation calling forth spirits and something akin to Satan himself appeared? That's (kind of) where we are with GPT.