As an analogy, Uber could crank up rates after the VC growth play was over to stoke revenue and profits because they have a duopoly with Lyft. LLM consumers can switch to Kimi models fairly trivially today, and whatever the frontier open model landscape looks like later. Model training and development is expensive, self hosted inference on open models not so much.
https://www.wheresyoured.at/the-openai-bubble/ has the math.
(a component of my work is currently building scaffolding so our organization can swap out commercial inference providers for on prem inference infra to derisk against the eventual rug pull when the math gets icky for LLM providers, while consuming as much subsidized tokens as we can until then, when it makes sense to use tokens for work)
Can you install a near-SOTA model on a cluster in a data center? Of course. Compliance and operations are the sticking points. I work in healthcare IT, and it's amazing how tight the data compliance requirements are. I can't have someone in Canada look at prod data. If we told hospitals that we were handing off PHI/PII to Chinese models, they'd end our relationship due to the long history China has of hacking Western networks and computers. They don't care how open and cheap things are.
Then, you have to keep up-to-date on the latest technology and right-size things in a very fluid market. If you sign a contract for hosting the model on a data center that's running what the SOTA is now in hardware, and someone comes through with a data center hardware or software product that makes that data center contract a disadvantage (maybe it's too expensive and the other party won't budge on the price), you might have to factor that into your offering's price, and that could put you at a disadvantage in your marketplace.
Google, MS, etc. all want to leverage the cloud model to make this be less of an issue for you, for a price. They have the ability to update you with the SOTA stuff in the data centers, because they're the ones driving that SOTA. They can say they host in the US and develop most of their stuff in the US.
Will that be enough of a moat?
Probably not for the levels of spending that are happening now, but over the long term, probably.
Customers can switch (although we can argue the speed and pain of doing so), and the speed at which they do will be a function of cost efficiency and demonstrable value (imho). A recent example of this is Broadcom and VMware [1], for example. When motivated, it can be done. If there is no objective, measured value being delivered, the spend will be cut. If the value delivered is measured, it will be enabled at a lower cost through cost optimization measures (ie self hosting) [2].
This is all to say: there is no moat, the revenue of inference providers is volatile and not assured in any measure. Caveat emptor.
[1] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
[2] Microsoft considers replacing ChatGPT and Claude with Kimi K3 to save $600M - https://news.ycombinator.com/item?id=49022984 - July 2026
If the second one were as easy as the first, I wouldn't have to be online at 9:00 to deploy stuff to prod tonight; the team in India would handle it. But customers write into the contracts that only US-based employees interact with prod systems. No amount of cajoling will get them to change their minds; they have data sovereignty, international telecommunications treaties, and HIPAA compliance to worry about. So I'll be pressing buttons tonight.
Could you swap out Anthropic or OpenAI or Google or whoever's models for Kimi? Yes. They're like other software these days, they're modular. What isn't modular is regulatory and geopolitical concern.
To setup a cross business kubernetes cluster will take 2 years with unknown results.
On Cloud, in Switzerland, you need to call Microsoft when you need new resources, so much for agility and minute infrastructure provisioning, and I heard the same for AWS.
> To setup a cross business kubernetes cluster will take 2 years with unknown results.
Do you seriously believe those times will not go down 95% if the CEO pushes for it to get done yesterday because it will save the company millions in expenses?
it has to be some amazing router and while the models are open-weights, the knowhow to run them efficiently surely is not?
Meanwhile the cost/benefit analysis doesn't move much even if you are paying 2x for tokens, and you don't need anything on prem.
Look at any computer in a big company. It isn't the fastest on the market, nor will it have the most RAM or largest monitor or fanciest keyboard. It is good enough at a good enough price point. Once it becomes possible and cheaper to host your own good enough open weight models, with all the benefits of keeping data internal to the company, then the big providers are cooked, so to speak.
Depends on the advantage it gives people and marketing of that advantage. You'd be surprised at how overpowered the average workplace laptop is. Each company I've been at has had at least some people who do non-technical roles using high-end hardware. Why? Because the account executive wants the fast machine and they get what they want.
You can apply the same to GenAI. Humans are notoriously bad at estimating actual needs when it comes to resource consumption. Best to have it and not need it than need it and not have it, especially if your competition just shelled out for SOTA.
And that's not even taking into consideration regulatory and customer concerns about where the AI you're serving requests with came from.
We have already seen tech workers at big name companies get whiplash from "leaderboards showing people using the most tokens!" as a good thing one month to being pressured to using fewer tokens a month later.
You'd be surprised. I've seen people request upgrades and get them. Not even execs, just employees. Even the average devices are probably overpowered these days. Chromebooks could do the vast majority of work in a corporate environment. Try getting workers to accept them, though.
Regardless, there are probably enough reasons to use proprietary Western AI for the time being, particularly in regulated industries, that poop won't completely hit the fan. Particularly if the CTOs start searching the phrase "Operation Aurora".
> We have already seen tech workers at big name companies get whiplash from "leaderboards showing people using the most tokens!" as a good thing one month to being pressured to using fewer tokens a month later.
Indeed. What do you do when you want a group of people to do something that they'd be otherwise adverse to doing at all? You rank them on it and let them get in a competition over who can do it the most. That's what tokenmaxxing was about: getting people to use the AI at all. Now that the people are using AI, we move onto another objective, which is getting them to use the tokens efficiently. The trick is measuring that efficiency.
Even if, hypothetically, Fable or a Fable-class model could seriously replace some headcount, it'll only gain further traction of it's actually cheaper than hiring humans. $50/MTok is expensive. Wouldn't be unreasonable to expect somewhere between ~$3k-$5k/month/developer in spend. Cheaper than a Junior in the HCoL areas (in the US), but not much cheaper in lower-to-average COL areas. Most acceleration will come from having the headcount + giving said headcount $3k-$5k/month in token budget, so now it just becomes a very expensive dev tool rather than a headcount replacement tool.
The idea that a $30k/year API bill will replace 2 $100k developers falls part outside of SFC/NYC. No CFO of a mid-market company in a LCOL area is signing off on $3k/month/dev API bills. They'll just hire juniors and cap their spend at $200/month.
(The claim felt so wild I wanted to check, and indeed, the private Google Cloud for the $125bn Australian pension fund was accidentally deleted by a provisioning misconfiguration. Any others?)
Turns out you can fuck up self hosting too.
I'd be with you if you claimed that the revenue hasn't translated into substantial profits. Being able to spend a lot of money to get less money back is not that impressive. But revenue by itself is on a dramatic rise as capabilities improve
That's not a meaningful number, and even if it were, quadrupled isn't nearly enough.
Completely false.
AI and AI related revenues are growing exponentially.
Or are you just adding nonsense about "yeah but yeah but no value"?
Please try and provide one for such strong claims.
If AI-related expenses are also growing exponentially, and they are growing exponentially faster, it doesn't matter that revenue is growing exponentially.
The AI funding has also now absolutely baked in exponential growth of expenses, because that's how debt works. A slow exponential, hopefully, but an exponential none-the-less.
Something Hacker News needs to be periodically reminded of is that we are the field getting the most out of AI, and it's not even close. That's great for us. But the stocks aren't priced for "a pretty nice coding tool". They're priced for every field in the world getting even more value out of this than our field is getting now. That is, frankly, not happening anywhere near fast enough for the spending and stock valuations. When you don't have all the engineering guardrails that are present in software engineering [1], suddenly the AI is, ahem, exponentially less useful.
As I say in that post, watch your AI actually doing something, even the frontier models. Watch the thinking traces. Watch how many times they bang into a guardrail of some sort; a failing test, a failing compile, a linter failure, a bash script that doesn't work, all those things. How much value would you get out of an AI coding assistant if the first time it banged into a guard rail it was done and you had to stop using it for that task? How much value would you get out of an AI coding assistant if instead it silently failed and just proceeded forward with errors that you lack the infrastructure to easily detect? In the first case, it would be fairly modest, almost certainly not worth the money, and in the second, it would be worth paying to not use.
Even in our field, while the rate of code output has increased substantially, the rate of value generation increase has been quite a bit more modest. I have observed, and heard from a number of other places, that while my own output has increased somewhat we still generally can't plan on being able to work with other teams at much faster a rate than we used to.
There's a viable business here but I can't see how all these companies expect to be returning all this revenue in any financially sensible period of time. They're all spending like if only they spend enough they can own about %900 of the market in three years. They can't all do that, even accounting for "AI makes the market bigger".
And they're wildly vulnerable to some new solution coming out that obsoletes all this spending, like an ASIC that starts running a popular model directly (especially if model capabilities plateau), meaning that all this nVidia GPU spending is so much dead silicon. Or someone comes out with a much more efficient way to train models. There has to be some insight we're missing; humans do not learn what they do by having the entire contents of the Internet poured through their head hundreds of times over. We are far more efficient with our training data. What if someone works out a solution to that and we don't need to spend billions on GPUs but only millions? The whole spending proposition could collapse overnight and the companies that suddenly have three orders of magnitude too much hardware and the debt to match would be up a creek without a paddle.
[1]: https://jerf.org/iri/post/2026/programming_is_engineering/
Just emphasizing that as, due to spending far too much time online the past week, I've been seeing a fair bit of this. "AI is definitely gaining popularity because all the software companies I know are going all in on it."