So yes, I think your scenario is likely to eventually happen, but there will be a much more powerful, capable frontier model then.
It's almost a given that whatever is frontier intelligence today will run on a potato in a few years.
I'm pointing out that there's no known information theoretic constraint about the impossibility of frontier AI models being improved to fit/run on a small GPU.
Please do not make up plausible sounding science facts.
If I were to say you could put a motorcycle in my car’s trunk, it would be perfect valid for me to say there are space constraints that make your idea unlikely. The same is true in this discussion, even though I have not computed the exact dimensions of the motorcycle and my car’s trunk.
Sure, there could be some point between a midrange consumer GPU and a pocket calculator where you can't fit enough 'intelligence'. But we really have no idea if the constraint is information theoretic or something completely different. Demonstrating that is the hard part, not finding the exact number of bits.
Talking about motorcycles in car trunks is just lazy false analogy here.
There are constraints of course- training takes way longer.
We don't really have the tools to reason about this stuff yet. Exciting times.
People will claim to have “enough” even though they already have the equivalent of last years capabilities locally.
Even if it did, it still doesn't make much economic sense running a model locally vs on a datacentre.
For example, I managed to just about squeeze a Q2 quant of Qwen 3.7 27b on my 9070XT. I get around 60tps decode (slightly faster prefill). _but_ it uses 300W of power to do so. At UK electricity rates of 30c/kWh this works out at something like 42c/MTok. I can get far far better models on openrouter cheaper than that, plus I'm not horrendously constrained on context length.
Like even if you run it in a datacenter in this scenario, you could do it on a cheap GPU instance in Azure, you still wouldnt need OpenAI or Anthropic specific clouds.
>uses 300W of power to do so.
There are plenty of people with phat electricity pipes in their on prem server rooms that have been vacated for cloud. Companies who want the benefits of AI but dont want the risk of sending their data to foreign API endpoints.
But Murphy's law is dead. No future chip will leapfrog easily current chips because we have reached hard phyical limits in chip density and downsizing. Huang's law by Jensen Huang focuses on something else and that is token performance per Watt at scale.
Blackwell needs double TDP than Hopper and Rubin again needs almost double TDP on a rack but in the end Rubin will be like 100x token performance per watt on a scaled data center. This means you have more energy need but you get multiples of token performance because you start scaling in the data center.
The local chip will never be able to keep up with the data center scaling economics. This is why everyone is so crazy about building data centers because they can see the economocs behind it.
What people don't seem to understand if tokens become more available and cheaper then not only more people can use them but a single person can use more as well. Why should you be limited to 1 AI agent? Why can't have you have multiple agents running on multiple devices daily for you?
This is why demand will grow exponentially with the growth of token economics. We have seen it for the last few years and much more is yet to come.
>The local chip will never be able to keep up with the data center scaling economics.
Assumes the software has been completely solved.
>Why should you be limited to 1 AI agent? Why can't have you have multiple agents running on multiple devices daily for you?
At some point we cap out the bandwidth of the human to keep up with their mistakes.
If it does happen then NVidia will sell a lot of 5070s though!
the biggest winner in that scenario would be ai providers, who suddenly have a capable model that they can serve much more efficiently. and the incumbents have a whole lot of compute. wouldn't anthropic and openAI just start offering that open weights model at prices that nobody else could compete with?
RTX 5070 prices go up ~N times. Nvidia makes more money because it's easier to make these things than it's to make a GB300.