It’s a wee bit early to call this. Let’s see what the top labs release in the next year or two, yeah?
GPT-4 was released only 15 months ago, which was about 3 years after GPT-3 was released.
These things don’t happen overnight, and many multi-year efforts are currently in the works, especially starting last year.
There are probably some situations where it suffices to use a small model, but for most purposes, I'd prefer to use the state of the art, and I'm eager for that state to progress a little more.
I'm guessing in the future, we'll see a lot more automatic inference "on our behalf", and you won't care or notice if its using Llama5:3b or whatever comes out then.
I'm betting that in a few years we'll see LLMs baked into a ton of stuff - lots of "simple" things mostly, like email summary/rewording, and similar "light touch" use cases. Maybe light multi-modal work like photo labeling. Probably expanded to be running in every other SaaS applications doing god knows what. For those, the small models would be more than enough, and much cheaper to run in large volumes.
I'm guessing "chat with a bot to ask questions" will be a small amount of the inference that happens on our behalf, but will use the valuable SOTA model use case.
Not only big tech is part of it but billion dollars startups are popping everywhere from China to US and Middle East.
At some point everyone will realize that it is becoming a commodity, and it is very expensive to train, then only those with wither a structural advantage to lower price (eg Google) or a true goal of being on the high-end/SOTA of the market (OpenAI, Anthropic) will keep going.
I'm finding it difficult to estimate the size of this workload compared to continued training of foundation models. Perhaps it depends on whether there are new architectural breakthroughs that require retraining of foundation models.
And what about non-language tasks such as interpreting video and 3D sensory data? This is potentially huge, but between huge peaks there is often a valley of unknowable depth and breadth.
Yea yes video probably requires a lot of GPUs to train. And a lot of source material to train against. And a use case. Which again, most companies don’t have.
Model development is clearly here to stay, and clearly valuable. Models from every other company, either foundation or fine tuned - I’m not sure that emperor is wearing many clothes any time soon.
next wave driving demand can be actual new products developed on LLMs. There are very few usecases currently well developed besides chatbots, but potential is very large.