Those things can all be done today on a $250 used video card and pennies of electricity
Heck, most large enterprise moved to usage based billing and are still happily paying for it. They are force multipliers for your top talent, and when a top engineer is being paid $500k a year, doubling their output for $500/day is a no brainer.
Take a $10 claude sub and fill it with Fable/Opus for a month (meaning use all your tokens) and then use one of the many token/session recording tools (I like agentsview).
It will show you how much your token cost is at current rates if using an api. I use the $20 plan and I don't use all my tokens, and am around 50x.
We're going to see Apple and Google compete over services and AI/OS integration instead, it will probably be years before your OEM takes local models seriously.
Running KIMI on a phone is not possible today and I agree with you that it will "probably be years before..." it is.
But how many years do you guess? I personally do not think it will take even 10 years for the situation to be commonplace.
Do you personally remember how far smartphones progressed in the past 10 years? It's not as long a time as you think it is, the limits of what a smartphone GPU is capable of did not substantially change in that time. Nor did the amount of onboard RAM that we include in the package. This is true even for Nvidia's ARM SOCs, frankly.
Apple, Microsoft and Google all eventually want to enforce OS-level lock-in for the most profitable AI services (eg. their own). It's much more attainable and profitable to use that lock-in to sell you exclusive service integration, the local AI revolution probably won't begin on their hardware.
IMO it won't be possible for the foreseeable future. There's essentially zero possibility that phones will gain the hardware capacity to run today's Kimi, so the only other alternative is to squeeze the power of today's Kimi into something that can fit on a smartphone, which also seems fairly unlikely considering the current rate of progress.
Phones have zero HBM today.
> There are famous laws of growth that put terabytes at just a few years away
I assume you're referring to Moore's law here, but if you are, it doesn't really apply to HBM in the same way, especially in a smartphone form factor where LPDDR is the only practical option due to heat and energy constraints, and a variety of architectural complexities specific to HBM that make a TB of it in a smart phone something far beyond what we can hope for in any timeline we can project today.