To me there is still so much left to do.
And even if we truly did multiple laps over all the data in the World there are still different new architectures and strategies left to try.
To me there is still so much left to do.
And even if we truly did multiple laps over all the data in the World there are still different new architectures and strategies left to try.
your comment really belies the desperation that exists now, these models are stuck where they are (hint it’s a natural limit), you are talking about exponentiation of cost to get what a 10% improvement? 5%? They are very few places for which it’s net positive to run them now, and most of those are incredibly shitty things like creating trash marketing content to drown us all in average inanity
I really feel bad for this next generation, they will just be constantly inundated with generated crap, so much of the high fidelity of conversation and meaning is and will be lost.
And I am not talking about predicting the future, but more predicting the next action to take based on current state, sensor data in a more seamless way. Like a human being reacting to different input, by moving their muscles etc. There would be huge amount of training data from there that could be incorporated into a single model.
Like self-driving cars?
Self-driving cars is an engineering problem, let alone an AI problem, and we still cannot solve it despite trillion dollar economic incentives.
Just putting together some LLMs on a fuckton of data does not work. Tesla tried that, and failed.
Completely different solution applied to a completely different problem with completely different risk and quality tolerances with completely different mitigations.
Given that LLMs are inaccurate around 5-10% of the time each step will compound the error rate until you are better off flipping a coin.
You can ask them to do math equation which takes steps and if they are trained in that for certain problems they are accurate near 100 percent of the time.
Like ask gpt-4o to solve different variations of
"""What is the answer to 2x + 7 = 31?"""
If the numbers are of similar magnitude and simplicity, it will follow the same steps and be right 99%+ times, and I'm only not saying 100%, because I haven't tried it enough, but I don't see it being wrong.
For example """What is the answer to 2x + 4 = -6?"""
Just run a test yourself. Do random integers within 0 - 20, it will definitely not be incorrect 5% - 10% time. It will be correct 99%+ time.
Where is this number 5% - 10% even coming from? You could also keep asking it "What is the capital of France?" and it's going to be right 99%+ of the time.
And the 5-10% is on average and gets significantly worse as you expand the context length which is also something you want for an agent.
Based on what you are attempting to do you could get any average in the end.
That happy discovery was never really a linear improvement path, though. We had an explosion of capability, but all along there have been active questions about how far the improvements would go with the current approach.
I think the point that a lot of researchers are making is that that we're starting to see those limits (with LLMs, at least).
There are also a lot of questions around business model and cost/value prop. Training and running these things at scale is enormously expensive. I'm seeing a lot of FOMO and gold rush mentality in the space, similar to the online streaming wars, and I'm not convinced of the long term viability of a lot of the companies. Especially once open models like llama are "good enough" and become commodities.
Of course, it's still early days and there's a ton of room for discovery, but it looks like we'll hit a limit with the current approach pretty soon.
Personally, I'd be OK with that. With the current state of things we have an interesting toy that can sometimes do useful work. It's an incremental quality of life improvement and another good tool in the chest, but it's not a civilization impacting technology.
That's probably for the best.
And only after 1.5 years? And especially of we just had an happy surprise like you mentioned. How does it make sense to already start claiming that we have hit the limits. How do we know there is no more scaling, optimisations and happy surprises?
> That happy discovery was never really a linear improvement path, though. We had an explosion of capability, but all along there have been active questions about how far the improvements would go with the current approach.
> I think the point that a lot of researchers are making is that that we're starting to see those limits (with LLMs, at least).
The kinds of limitations we're "starting to see" are largely the same as they were a year ago. People were talking about it on here back then, but now it's becoming more apparent to more people as they get used to LLMs.
For those who saw it back then, this does look like we're hitting a limit. For others, not so much.
How do active questions about a technology imply we are approaching a brick wall?
How could researchers without having access to the latest state of the art - by OpenAI or any other unknown companies be able to even test that we could be approaching a brick wall? It seems to me that it would take trillions to find out what the exact limit is.
It's possible that we will get diminishing returns, but I don't see how we can confidently claim or know it?
> The kinds of limitations we're "starting to see" are largely the same as they were a year ago. People were talking about it on here back then, but now it's becoming more apparent to more people as they get used to LLMs.
I don't follow. GPT-3.5 was borderline useless at reasoning. But it still seemed amazing and what I wouldn't have thought to be possible in any near future.
And then GPT-4 was a crazy advancement over that to me. And I've been using it daily since it was available, for various use-cases. Are you saying we are seeing the limitations of GPT-4 specifically? Because, sure, GPT-4 is far from AGI, but I don't see how this implies that further scaling, optimisation, training data improvements, techniques like multi modality and other potential strategies that I might not be aware of couldn't bring another explosive step?
Also the fact that GPT-4 reasoning skill hasn't been reproduced by anyone else so far seems to leave me thinking that everyone except OpenAI are clueless. Claude Opus is close, like I've said before, but not quite GPT-4 levels in specific reasoning tasks that I'm using the API for.
If you can't reproduce GPT-4, how could we trust the assessment that we have hit a limit?
Not sure what that means. Why are you marking those as "quotes".
The last actions brought so many returns. And it's unknown what the exact effect would be in adding more modalities, training data, optimisations and even just plain parameters.
Text as training data can only get you so far. Giving real time sensory data from many fields could allow LLM like system to control robots and get even more data from real life. E.g. robot hand movements, object tracking data, all of that to be fed into LLMs, and see how it would work.
- something that can't be modeled because there's no training data
- something that can't be modeled because it's fundamentally stochastic
- something that can't be modeled because the discrepancy in simulating the generating process, for your specific model, can, basically, be made arbitrarily large
I think a very common error when it comes to personal learning or progress is confusing a plateau with a brick wall. The reason is, unless you have already walked the path, it’s not possible to differentiate them. And when it comes to progress, no one has already walked the path, hence no one knows actually.
What if it already did, and it's called GPT4-o? Like sure, OAI realized it was a mostly marginal improvement over 4 after it was finished. But did they know that ahead of training?
I would have expected gpt-5 at least initially to be much, much slower, and they have only recently talked about starting on it.