I predict we'll get a few research breakthroughs in the next few years that will make articles like this seem ridiculous.
I predict we'll get a few research breakthroughs in the next few years that will make articles like this seem ridiculous.
Re training data - We have synthetic data, and we probably haven't hit a wall. Gpt-5 was only 3.5 months after o3. People are reading too much into the tea leaves here. We don't have visibility into the cost of Gpt-5 relative to o3. If it's 20% cheaper, that's the opposite of a wall, that's exponential like improvement. We don't have visibility into the IMO/IOI medal winning models. All I see are people curve fitting onto very limited information.
A "frozen mind" feels like something not unlike a book - useful, but only with a smart enough "human user", and even so be progressively less useful as time passes.
>Doesn't seem like a problem that needs to be solved on the critical path to AGI.
It definitely is one. I know we are running into definitions, but being able to form novel behavior patterns based on experience is pretty much the essence of what intelligence is. That doesn't necessary mean that a "frozen mind" will be useless, but it would certainly not qualify as AGI.
>We don't have visibility into the IMO/IOI medal winning models.
There are lies, damn lies and LLM benchmarks. IMO/IOI is not necessarily indicative of any useful tasks.
But every time you tried to get him to do something you'd have to teach him from first principles. Good luck getting ChatStein to interact with the internet, to write code or design a modern airplane. Even in physics, he'd be using antiquated methods and assumptions, this getting worse as time progresses(like sib comment I believe was alluding to).
And don't even get me started on the language barrier.
I recently read this short story[1] on the topic so it's fresh on my mind.
I‘ve yet to see a convincing article for artificial training data.
This. Lack of any way to incorporate previous experience seems like the main problem. Humans are often confidently wrong as well - and avoiding being confidently wrong is actually something one must learn rather than an innate capability. But humans wouldn't repeat same mistake indefinitely.
The feedback you get is incredibly entangled, and disentangling it to get at the signals that would be beneficial for training is nowhere near a solved task.
Even OpenAI has managed to fuck up there - by accidentally training 4o to be a fully bootlickmaxxed synthetic sycophant. Then they struggled to fix that for a while, and only made good progress at that with GPT-5.
But I agree that being confidently wrong is not the only thing they can't do. Programming, great, maths, apparently great nowadays, since Google and OpenAI have something that could solve most problems on the IMO, even if the models we get to see probably aren't models that can do this, but LLMs produce crazy output when asked to produce stories, they produce crazy output when given too long confusing contexts and have some other problems of that sort.
I think much of it is solvable. I certainly have ideas about how it can be done.
I think the next iteration of LLM is going to be "interesting", i.e. now that all the websites they used to freely scrape have been increasingly putting up walls.
Except nvidia perhaps
Uber successfully turned a war chest into a partial monopoly of all ride hailing for significant chunks of the world. That was the clear plan from the start, and was always meant to own the market so they could extract whatever rent they want.
Amazon reinvested heavily while competition floundered in order to literally own the market, and has spent every second since squeezing the market, their partners, everyone in the chain to extract ever more money from it.
None of those are even close to buying absurdly overpriced hardware from a monopoly and reselling access to that hardware for less than it costs to run and doing huge PR sweeps about how what you are building might kill everyone so we should obviously give them trillions in government dollars because if an American company isn't the one to kill everyone than we have failed.
"Not terribly knowledgeable in the field"?
You’re right in that it’s obviously not the only problem.
But without solving this seems like no matter how good the models get it’ll never be enough.
Or, yes, the biggest research breakthrough we need is reliable calibrated confidence. And that’ll allow existing models as they are to become spectacularly more useful.
Ha, that almost seems like an oxymoron. The previous encounters can be the new training data!
What would be the point of training an LLM on bot answers to human questions? This is only useful if you want to get an LLM that behaves like an already existing LLm
But memory is a minor thing. Talking to a knowledgeable librarian or professor you never met is the level we essentially need to get it to for this stuff to take off.
And now, in some cases for a while, it is training on its own slop.