Second off, this doesn’t work from a power consumption standpoint. When I run qwen3.6-35b, a far smaller model than op is suggesting, power usage spikes to 150-200W during inference. To fit a 1T model in the palm of my hand, the amount of processing required doesn’t fit the amount of power available.
Now I’m not saying this will never happen - there are some great leads, e.g. burning models directly on to a chip - but op’s scenario is definitely not happening in two years. Maybe 5, a lot more likely 10, unless of course local ai is made illegal
> Datacenters are willing to pay $50k for a single high end GPU.
its true for now, because capital is flowing like a torrent, but how long will that last if returns start to be expected (aka the bubble pops)?That doesn’t change until production capacity exceeds the datacenter demand. When that happens, they’ll start selling them down the market until it eventually reaches phones and toasters and whatever. But not in two years.
* Software inference optimizations
* Heavy quantization
* Chips with hardcoded transformer architecture
* Much cheaper HBM
* Much sparser models - 1T total with ~1-10B active params e.g.
* Not to mention - 2 years of today's frontier models writing RTL and kernels at superhuman levels.
Absolutely. I'd be surprised if they couldn't 2x performance in the next year. Still doesn't make a 1T model fit on your phone.
> * Heavy quantization
I think this is a dead end if you're trying to fit a 1T model into a phone. Makes much more sense to train a model that's designed to be small, than train a model that's smart and then quantize it into stupidity.
> * Chips with hardcoded transformer architecture
Totally, this will probably work great. Now good luck booking fab time any time in the next 2 years.
> * Much cheaper HBM
Totally, this will probably work great. Now good luck booking fab time any time in the next two years.
> * Much sparser models - 1T total with ~1-10B active params e.g.
Fewer active params helps with the speed of token generation, but if the whole model doesn't fit into ram it doesn't solve the issue of having to constantly stream portions of the model from disk to ram.
> * Not to mention - 2 years of today's frontier models writing RTL and kernels at superhuman levels.
IMO this is a delusional myth-making idea being sold to us by ai companies. Machines that generate output based on statistical averages won't generate genuinely new ideas. They can help us try out ideas faster, but they're simply not capable of the kind of creativity and understanding required to push a field forward, except incrementally.
I absolutely agree that models are going to advance on to “edge” hardware over the next few years by becoming small + specialized.