I'm very bad at using power drills.
I fear on-prem AI is likely to become as popular as on-prem servers without Cloudflare using self-hosted email are today: that is to say, people have heard of it but the skillset is almost popularly eviscerated, external policies make it progressively impractical, and anyone who does it is 'niche'. While basic guides will exist, obtaining top-level output will probably require many moons of concerted effort.
Basically: AI is SaaS for thinking.
Consider another perspective: They don't have to keep up. Once a model is good enough for a task, the model can stay. A hammer is a hammer and a hammer from 100 years ago still has most of its utility.
Similar to the hammer, it's not unreasonable to think that some classes of work will simply be solved by some model generation and whatever happens at the frontier after that does not matter all that much for work that puny humans do.
Then, of course, there will be a time where all of this is moot: Absolutely no human will want a human to diagnose their medical issues. That is not a skill deterioration issue. We simply will concede that we are not able to do it as well as a more capable system can, and without much fanfare, increasingly delegate, as we have always done.
Even if this occurs, and I don't trust well-resourced humans to allow their existing apex-predator positions in present era capitalism to be overturned, the action - as far as either humanity or AI is concerned - will still be at the forefront of possibility: a front by definition invisible to old models. And someone has to pay for the hardware to be there. Do we (a) allow private-sector dominance, effectively depowering traditional nation states and empowering a private cabal beyond historically conceivable levels (b) nationalize thought (c) head in sand and pretend it will all go away?
Most of the world seems to be with strategy C right now, strategy A is the advancing default and has already achieved extra-terrestrial reach with a threat of extra-terrestrial persistence, and strategy B is potentially scarier than the other outcomes if it goes wrong but might be lovely, if you believe in nordic state funds, solarpunk futures and socialist utopia.
Interesting times. By the way, if anyone with AI capitalization reads this, I'm looking for investment to feed humans more efficiently and have a NASDAQ reverse merger under negotiation and effectively priced out with board buy in. Just need capital support. https://infinite-food.com/
Concentration of power exists when the model makers are the same as (or control) the inference providers. Making a model is capital intensive, so there aren't many of them. Providing inference is not: I don't even need to own GPUs; I can rent them from those who do and then sell by the token. B300s cost less than $4 an hour currently.
Cloud can even be more effective at lowering concentration of power than on premise. Asking people to individually buy $20,000 of compute equipment plus power and cooling equipment to run a frontier model is not something they're going to do if they can just pay four-tenths of a cent per output token. If the only cloud inference providers are the big proprietary US titans, that means you're going to get far more power concentration than if open source inference providers are an alternative, because then I can just switch my API endpoint.
Edit: Mis remembered the timeline I saw not 2 weeks, 3 months, but still I think my point stands.
https://x.com/yaroslavvb/status/2067367657272422584 https://x.com/voratiq/status/2067667800643268928 https://arena.ai/leaderboard/agent
https://news.ycombinator.com/item?id=48567759
Commenters there were saying GLM 5.2 was roughly equivalent to Opus 4.8 in coding prowess, based on personal experience of the people commenting. Opus 4.8 came out on May 28 this year (so more like 3 weeks ago), GLM 5.2 came out 2 days ago.
After a while, you do start to start to skip a couple rounds of open source models until there's a notable release. That, and the resources needed to run them are increasingly bought up by the owners of frontier models
LLM output is unreliable, so we still need to judge it. If I want to be able to judge code, I must have worked with it to a certain extent. So the unreliable tool does not help me much if I don't want to accept the unreliability.