Yes. This is exactly why I'm skeptical of AI doomerism/saviorism.
Too many people have been looking at the pace of LLM development over the last two (2) years, modeled it as an exponential growth function, and come to the conclusion that AGI is inevitable in the next ${1-5} years and we're headed for ${(dys|u)topia}.
But all that assumes that we can extrapolate a pattern of long-term exponential growth from less than two years of data. It's simply not possible to project in that way, and we're already seeing that OpenAI has pivoted from improving on GPT-4's benchmarks to reducing cost, while competitors (including free ones) catch up.
All the evidence suggests we've been slowing the rate of growth in capabilities of SOTA LLMs for at least the past year, which means predictions based on exponential growth all need to be reevaluated.
If you have a machine that converts mass into energy and then uses that energy to increase the rate at which it operates, you could rightfully say that it will level off well before consuming all of the mass in the universe. You just can't say that next week after it has consumed all of the mass of Earth.
Example:
w której gwarze jest słowo ekspres i co znaczy?
Słowo "ekspres" występuje w gwarze śląskiej i oznacza tam ekspres do kawy. Jest to skrót od nazwy "ekspres do kawy", czyli urządzenia służącego do szybkiego przygotowania kawy.
The correct answer is that "ekspres" is a zipper in Łódź dialect.But we could add internal thoughts-- we could make the model generate tokens that aren't part of its output but are there for it to better figure out its next token. This was tried QuietSTAR.
Hochreiter is also active with alternative models, and there's all the microchip design companies, Groq, Etched, etc. trying to speed up models and reduce model running cost.
Therefore, I think there's room for very great improvements. They may not come right away, but there are so many obvious paths to improve things that I think it's unreasonable to think progress has stalled. Also, presumably GPT-5 isn't far away.
Why do we presume that? People were saying this right before 4o and then what came out was not 5 but instead a major improvement on cost for 4.
Is there any specific reason to believe OpenAI has a model coming soon that will be a major step up in capabilities?
I assume that this won't take forever, but will be done this year. A couple of months, not more.
It feels like there’s an assumption in the community that this will be almost trivial.
I suspect it will be one of the hardest tasks humanity has ever endeavoured. I’m guessing it has already been tried many times in internal development.
I suspect if you start creating a feedback loop with these models they will tend to become very unstable very fast. We already see with these more linear LLMs that they can be extremely sensitive to the values of parameters like the temperature settings, and can go “crazy” fairly easily.
With feedback loops it could become much harder to prevent these AIs from spinning out of control. And no I don’t mean in the “become an evil paperclip maximiser” kind of way. Just plain unproductive insanity.
I think I can summarise my vision of the future in one sentence: AI psychologists will become a huge profession, and it will be just as difficult and nebulous as being a human psychologist.
I'm in the process of spinning out one of these tools into a product: they do not. They become smarter at the price of burning GPU cycles like there's no tomorrow.
I'd go as far as saying we've solved AGI, it's just that the energy budget is larger than the energy budget of the planet currently.
High temperature will obviously lead to randomness, that's what it, evening out the probabilities of the possibilities for the next token. So obviously a high temperature will make them 'crazy' and low temperature will lead to deterministic output. People have come up with lots of ideas about sampling, but this isn't really an instability of transformer models.
It's a problem with any model outputing probabilities for different alternative tokens.
What if they have two teams? One dedicated to optimizing (cost, speed, etc) the current model and a different team working on the next frontier model? I don't think we know the growth curve until we see gpt5.
I'm prepared to be wrong, but I think that the fact that we still haven't seen GPT-5 or even had a proper teaser for it 16 months after GPT-4 is evidence that the growth curve is slowing. The teasers that the media assumed were for GPT-5 seem to have actually been for GPT-4o [0]:
> Lex Fridman(01:06:13) So when is GPT-5 coming out again?
> Sam Altman(01:06:15) I don’t know. That’s the honest answer.
> Lex Fridman(01:06:18) Oh, that’s the honest answer. Blink twice if it’s this year.
> Sam Altman(01:06:30) We will release an amazing new model this year. I don’t know what we’ll call it.
> Lex Fridman(01:06:36) So that goes to the question of, what’s the way we release this thing?
> Sam Altman(01:06:41) We’ll release in the coming months many different things. I think that’d be very cool. I think before we talk about a GPT-5-like model called that, or not called that, or a little bit worse or a little bit better than what you’d expect from a GPT-5, I think we have a lot of other important things to release first.
Note that last response. That's not the sound of a CEO who has an amazing v5 of their product lined up, that's the sound of a CEO who's trying to figure out how to brand the model that they're working on that will be cheaper but not substantially better.
[0] https://arstechnica.com/information-technology/2024/03/opena...
The current crop of benchmarks might not reflect these gains, by the way.
They are by no means bad, but I am now mostly interested in long context competency. We need benchmarks that force the LLM to complete multiple tasks simultaneously in one super long session.
Anyway here's the sales page. the widget subscription is so premium you won't even miss the subscription fee.
We should still be skeptical because often want to claim to be better or have unearned answers, but I don't think the motive to lie is quite as strong as a salesman's.
It's not peer-reviewable in any shape or form.
That said, doing that is slow and people will need to make decisions before that is done.
It's weird. In the late 2010s it seems like people were wising up to the idea that you can't implicitly trust big tech companies, even if they have nap pods in the office and have their first day employees wear funny hats. Then ChatGPT lands and everyone is back to fully trusting these companies when they say they are mere months from turning the world upside down with their AI, which they say every month for the last 12-24 months.
In the 2000s we only had Microsoft, and none of us were confused as to whether to trust Bill Gates or not...
https://www.wheresyoured.at/pop-culture/
> What makes this interview – and really, this paper — so remarkable is how thoroughly and aggressively it attacks every bit of marketing collateral the AI movement has. Acemoglu specifically questions the belief that AI models will simply get more powerful as we throw more data and GPU capacity at them, and specifically ask a question: what does it mean to "double AI's capabilities"? How does that actually make something like, say, a customer service rep better? And this is a specific problem with the AI fantasists' spiel. They heavily rely on the idea that not only will these large language models (LLMs) get more powerful, but that getting more powerful will somehow grant it the power to do...something. As Acemoglu says, "what does it mean to double AI's capabilities?"
I’m about Zucks age, and have been following his career/impact since college; it’s been roughly a cosine graph of doing good or evil over time :) I think we’re at 2pi by now, and if you are correct maybe it hockey-sticks up and to the right. I hope so.
If LLMs end up being the platform of the future, Zuck doesn't want OpenAI/Microsoft to be able to monopolize it.
> Other companies sell widgets. We have a bunch of widget-making machines and so we released a whole bunch of free widgets. We noticed that the widgets got better the more we made and expect widgets to become even better in future. Anyway here's the free download.
Given that Meta isn't actually selling their models?
Your response might make sense if it were to something OpenAI or Anthropic said, but as is I can't say I follow the analogy.
- flex
- deal a blow to Altmann
But it has nothing to do with LLMs (and interestingly enough they aren't opening their recommendation tech).
I mean, going by their own model evals on various benchmarks (https://llama.meta.com/), Llama 405b scores anywhere from a few points to almost 10 points more than than Llama 70b even though the former has ~5.5x more params. As far as scale in concerned, the relationship isn't even linear.
Which in most cases makes sense, you obviously can't get a 200% on these benchmarks, so if the smaller model is already at ~95% or whatever then there isn't much room for improvement. There is, however, the GPQA benchmark. Whereas Llama 70b scores ~47%, Llama 405b only scores ~51%. That's not a huge improvement despite the significant difference in size.
Most likely, we're going to see improvements in small model performance by way of better data. Otherwise though, I fail to see how we're supposed to get significantly better model performance by way of scale when the relationship between model size and benchmark scores is nowhere near linear. I really wish someone who's team "scale is all you need" could help me see what I'm missing.
And of course we might find some breakthrough that enables actual reasoning in models or whatever, but I find that purely speculative at this point, anything but inevitable.
The problem with this strategy is that it's really tough to compete with open models in this space over the long run.
If you look at OpenAI's homepage right now they're trying to promote "ChatGPT on your desktop", so it's clear even they realize that most people are looking for a local product. But once again this is a problem for them because open models run locally are always going to offer more in terms of privacy and features.
In order for proprietary models served through an API to compete long term they need to offer significant performance improvements over open/local offerings, but that gap has been perpetually shrinking.
On an M3 macbook pro you can run open models easily for free that perform close enough to OpenAI that I can use them as my primary LLM for effectively free with complete privacy and lots of room for improvement if I want to dive into the details. Ollama today is pretty much easier to install than just logging into ChatGPT and the performance feels a bit more responsive for most tasks. If I'm doing a serious LLM project I most certainly won't use proprietary models because the control I have over the model is too limited.
At this point I have completely stopped using proprietary LLMs despite working with LLMs everyday. Honestly can't understand any serious software engineer who wouldn't use open models (again the control and tooling provided is just so much better), and for less technical users it's getting easier and easier to just run open models locally.
OpenAI did a good move with making GPTo mini so dirty cheap that it's faster and cheaper to run than LLama 3.1 70B. Most consumers will interact with LLM via some apps using LLM API, Web Panel on desktop or native mobile app for the same reason most people use GMail etc. instead of native email client. Setting up IMAP, POP etc is for most people out of reach the same like installing Ollama + Docker + OpenWebUI
App developers are not gonna bet on local LLM only as long they are not mainstream and preinstalled on 50%+ devices.
In my opinion, they've found that intelligence with current architecture is actually an S-curve and not an exponential, so trying to make progress in other directions: UX and EQ.
https://nicholascharriere.com/blog/thoughts-openai-spring-re...
[0] not technically complete depreciation, since for example 4o mini is widely believed to be a distillation of 4o, so 4o's investment still carries over into 4o mini
I don't know if the utility is worse than an LLM that is SOTA for 2 months that no one even bothers switching to however - at least the marvel slop is being used for entertainment by someone. I think the market is definitely prioritizing the LLM researcher over Disney's latest slop sequel though so whoever made that comparison can rest easy, because we'll find out.
I thought that was the allure, something that's camp funny and an easy watch.
I have only watched a few of them so I am not fully familiar?
Has there been any indication that we're improving the lives of millions of people?
If I need to babysit a junior developer fresh out of school and review every single line of code it spits out, I can find them elsewhere
For example, has anyone ever attempted image -> html/css model? Seems like it be great if I can draw something on a piece of paper and have it generate a website view for me.
(May work poorly of course, and the sample I think I saw a year ago may well be cherry picked)
Have you tried upload the image to a LLM with vision capabilities like GPT-4o or Claude 3.5 Sonnet?
There are already companies selling services where they generate entire frontend applications from vague natural language inputs.
My hope now is that someone will figure out a way to separate intelligence from knowledge - i.e. train a model that knows how to interpret the wights of other models - so that training new intelligent models wouldn't require training them on a petabyte of data every run.
I had a discussion with a friend about doing this, but for CNC code. The answer was that a model trained on a narrow data set underperforms one trained on a large data set and then fine tuned with the narrow one.
I think GPT5 will tell if OpenAI hit a plateau.
Sam Altman has been quoted as claiming "GPT-3 had the intelligence of a toddler, GPT-4 was more similar to a smart high-schooler, and that the next generation will look to have PhD-level intelligence (in certain tasks)"
Notice the high degree of upselling based on vague claims of performance, and the fact that the jump from highschooler to PhD can very well be far less impressive than the jump from toddler to high schooler. In addition, notice the use of weasel words to frame expectations regarding "the next generation" to limit these gains to corner cases.
There's some degree of salesmanship in the way these models are presented, but even between the hyperboles you don't see claims of transformative changes.
buddy every few weeks one of these bozos is telling us their product is literally going to eclipse humanity and we should all start fearing the inevitable great collapse.
It's like how no one owns a car anymore because of ai driving and I don't have to tell you about the great bank disaster of 2019, when we all had to accept that fiat currency is over.
You've got to be a particular kind of unfortunate to believe it when sam altman says literally anything.
that's interesting. Do you have a rough percentage of this?
Does this mean these connections have no influence at all on output?
If you are sitting on 1 billions $ of GPU capex, what's $50 million in energy/training cost for another incremental run that may beat the leaderboard?
Over the last few years the market has placed its bets that this stuff will make gobs of money somehow. We're all not sure how. They're probably thinking -- it's likely that whoever has a few % is going to sweep and take most of this hypothetical value. What's another few million, especially if you already have the GPUs?
I think you're right -- we are towards the right end of the sigmoid. And with no "killer app" in sight. It is great for all of us that they have created all this value, because I don't think anyone will be able to capture it. They certainly haven't yet.
e.g. Almost the entire market relies upon Attention Is All You Need paper detailing transformers, and it would be an entirely different market if Google had held that as a trade secret.
I really am not expecting Apple or Microsoft to discover AGI and ferret it away for profitability purposes. Strictly speaking, I don't think superhuman intelligence even exists in the domain of text generation.
Of course, I agree that Stability AI made Stable Diffusion freely available and they're worth orders of magnitude less than OpenAI. To the point they're struggling to keep the lights on.
But it doesn't necessarily make that much difference whether you openly share the inner technical details. When you've got a motivated and well financed competitor, merely demonstrating a given feature is possible, showing the output and performance and price, might be enough.
If OpenAI adds a feature, who's to say Google and Facebook can't match it even though they can't access the code?
Given how good the model is terms of the quality vs speed tradeoff, they must have something.
I would guess that in that timeline, Google would never have been able to learn about the incredible capabilities of transformer models outside of translation, at least not until much later.
Progress is not slowing down but it gets harder to quantify.
Frankly I just don't understand the economics of training a foundation model. I'd rather own an airline. At least I can get a few years out of the capital investment of a plane.
I’m also optimistic about building better (rather than bigger) datasets to train on.
I think you're just seeing the "make it work" stage of the combo "first make it work, then make it fast".
Time to market is critical, as you can attest by the fact you framed the situation as "on par with GPT-4o and Claude Opus". You're seeing huge investments because being the first to get a working model stands to benefit greatly. You can only assess models that exist, and for that you need to train them at a huge computational cost.
It feels like ChatGPT won the time to market war already.
I'll switch again soon as something better is out.
You may be right with the average person on the street, but I wonder how many have lost interest in LLM usage and cancelled their GPT plus sub.
This is similar to the early days of the search engine wars, the browser wars, and other categories where a user can easily adopt, switch between and use multiple. It's not like the cellphone OS/hardware war, PC war and database war where (most) users can only adopt one platform at a time and/or there's a heavy platform investment.