Yes, tinkerers and enthusiasts will continue to make use of them, but frontier companies will maintain near total dominance because they will be the only ones with access to the hardware.
Yes, tinkerers and enthusiasts will continue to make use of them, but frontier companies will maintain near total dominance because they will be the only ones with access to the hardware.
We're going to see Apple and Google compete over services and AI/OS integration instead, it will probably be years before your OEM takes local models seriously.
Running KIMI on a phone is not possible today and I agree with you that it will "probably be years before..." it is.
But how many years do you guess? I personally do not think it will take even 10 years for the situation to be commonplace.
Do you personally remember how far smartphones progressed in the past 10 years? It's not as long a time as you think it is, the limits of what a smartphone GPU is capable of did not substantially change in that time. Nor did the amount of onboard RAM that we include in the package. This is true even for Nvidia's ARM SOCs, frankly.
Apple, Microsoft and Google all eventually want to enforce OS-level lock-in for the most profitable AI services (eg. their own). It's much more attainable and profitable to use that lock-in to sell you exclusive service integration, the local AI revolution probably won't begin on their hardware.
IMO it won't be possible for the foreseeable future. There's essentially zero possibility that phones will gain the hardware capacity to run today's Kimi, so the only other alternative is to squeeze the power of today's Kimi into something that can fit on a smartphone, which also seems fairly unlikely considering the current rate of progress.
Phones have zero HBM today.
> There are famous laws of growth that put terabytes at just a few years away
I assume you're referring to Moore's law here, but if you are, it doesn't really apply to HBM in the same way, especially in a smartphone form factor where LPDDR is the only practical option due to heat and energy constraints, and a variety of architectural complexities specific to HBM that make a TB of it in a smart phone something far beyond what we can hope for in any timeline we can project today.
Those things can all be done today on a $250 used video card and pennies of electricity
Heck, most large enterprise moved to usage based billing and are still happily paying for it. They are force multipliers for your top talent, and when a top engineer is being paid $500k a year, doubling their output for $500/day is a no brainer.
Take a $10 claude sub and fill it with Fable/Opus for a month (meaning use all your tokens) and then use one of the many token/session recording tools (I like agentsview).
It will show you how much your token cost is at current rates if using an api. I use the $20 plan and I don't use all my tokens, and am around 50x.
I don't think we'll see home users being able to match even the low end clouds for a long time.
Longer term I think we'll see these uses of AI cluster into a few groups:
- maximal code / reasoning quality, at high prices (Fable)
- typical code / agents (sub-Opus, Terra)
- cheap but decent enough quality (think Deepseek / GLM / Luna)
- so cheap I don't care about utilization (Deepseek, and friends)
And also more niche ones:
- ultra fast with high quality answers (typically sub-SOTA). Cerebras / dedicated silicon type approaches, expensive.
- ultra fast with mostly-adequate answers, and an openness to retries, moving up to better models
I think the open models will dominate (not with individuals, but low cost providers) all except the top 1-2 of those categories, and there will be a continuous erosion on the big player's moats. The top categories are also where all the money is, but I'm not sure it can justify those investments long-term. I also think they will have to squeeze more money out of them to justify the investments, which will also drive people down the list.
Edit: clarifications.
But they're not. Meta, SpaceX, Microsoft, Amazon, they're all leasing out capacity to others. If they were truly constrained, we wouldn't see that happening.
I wouldn't rule out the possibility completely, but it won't be very common.