You are completely missing the bet these companies are making.
They think can outlast their competitors and capture a larger portion of the pie while the cost of inference keeps going down dramatically.
If you haven't been paying attention, the cost is about 1/100th of what it was in 2024. This is the trajectory pretty much every technology has followed.
Of course there will be market crashes and corrections and things like that and most companies won't survive, but the bet is that whoever survives ends up doing pretty well.
The only relevant number is the price to serve a frontier or near-frontier model.
Also, lost revenue from other services being degraded by shifting resources to supporting training/serving models (Google Search...)?
Because everyone is buying as they want to run their own models and not pay for a cloud service?
If everyone's running local then why are these larger companies dumping cash into data centres?
You need a cluster of 8-12 H100s to run the largest models locally.
It doesn't make sense to run these locally yet unless your use case also involves making it available for several dozen concurrent users.
I've seen people happily use AI that takes several minutes to generate text or edit an image because to them they already aren't using their computer when they tell it to start; they just grab their phone and walk away and come back only to check in on it.
I feel like people here and on other technology discussions -- although it's worse here -- don't seem to parse what being the minority means.
They know they're one of the few to have access to such incredible hardware -- whether it be rented or purchased for way too much cash -- but they only see their own kin; their own ilk. They only compare themselves to the best.
The reality is that nobody expects data centre speed nor power in their own home and are satisfied to just go "haha its thinking" and let their computer quietly tick in the background as opposed to paying outragious prices for subscriptions or hardware.
I'm editing photos with 32GB of DDR4 RAM at 3200mhz alongside a 3050 with a nearly decade-old mobo and ryzen 5800XT and at most? It takes 30 seconds and that's allowing the gpu to use my actual RAM as 'fallback' memory.
There are photoshop filters -now- that take longer than that before the LLM craze even began. Hell; I can train a lora in an hour or two.
This obsession with "more specs more data faster and faster" isn't going to win; it already isn't winning.
Deal with it or get wrecked.
They promise updates.
Who's running local? Image generation can make sense to run locally, but frontier LLM make no sense to run on your own hardware.
We are also within an arms race of training newer larger models with more speed while discontinuing older models.
Gemini/Chatgpt have already discontinued their models from 2024 (iirc) because they are using all their compute in serving/training newer models. Being quite frank, nobody is serving a model from 2024 as the intended use-case while having very little moat as open source models are catching up.
> Of course there will be market crashes and corrections and things like that and most companies won't survive, but the bet is that whoever survives ends up doing pretty well.
How so, by raising the prices? because the current prices aren't sustainable and I feel as if there would certainly be companies which will try for one reason or other to be cheaper to capture the market share because of the larger promise of whoever is able to get as market share. I had once thought about it and I don't think that even in an ideal world, they would end up doing pretty well given no moat.
Also even if a company survives and ends up being one of the survivors and makes profit in the ideal scenario you mention, then within some years other companies will try again and construct more datacenters and end up driving the prices down for everyone, so nobody knows how things might look down for 2-3 years let alone a decade, so I remain a bit skeptic currently so.
I had actually thought some on the economics of datacenters and I found it to be very related to power. The only ones which seems to be making money might be the power generators actually because power is the actual bottleneck rather than GPU's in datacenters from my understanding.
Though the power is raised at the cost of electricity bill increases for everybody including people living in houses. The job prospects are minimal as well, as a nation, aside from just getting investment just for the sake of it because AI's trendy right now, I feel like its a net negative deal for people living there.
If cost of inference goes down 100x, would need 100x more demand. This makes overspending on GPU even more irrational. Jevons this, Jevons that but ultimately irrelevant. At end of day, leading players, hungergame winner candidates is saddling themselves with so much debt, even if they survive, post crash they are immediately uncompetitive against new entrant with blank slate and newer gen, more efficient GPUs that will be cheaper to buy/operate post crash when hardware prices will revert to mean.
It doesn't matter if some of the current players survive, they've basically stabbed and weakened themselves so much any healthy upstart in the future can wipe them out unless they lock in legislative protection... safety regulations, ban open source models etc.
That is the new bet, regulatory capture moat, because economic bet is entirely lost, especially with open models eroding mote.
A company losing out on a risky bet and failing, letting a new upstart rise up is pretty natural. A large fraction of experts from the failed companies continue at the new ones, business as usual. There are engineering teams at $BIGTECH now full of OS, database, or compiler experts from XP, Sun, HP etc.
This natural ability of companies to take risky bets is what made silicon valley successful.
It's independent of the former, but not the later. That is my point.
Businesses and business models fail all the time, does not mean the 'broader economy is going to explode'.
It could, sure. But that has been predicted several hundred times and happened only a few times.
Business models fail all the time, trillion business models that underpins entire macro economic growth engine exploding and causing massive contagion happens rarely, not even generationally. But when they do the leading bubble mathematic indicators are more clear.
Over 14 years, Uber burnt ~31 Billion dollars and in nearly whereas the amount invested within AI seems to be within Trillions at this point and in near future with a product which doesn't have much moat and shaky financials on profit.
It worked out for Uber, they are wildly profitable now after spending a decade losing money.
Also worked out as Amazon managed to outlast all the dotcom era e-commerce competitors while being unprofitable.
Uber have a 10% margin, which is definitely not what I'd consider wildly profitable. (Their post tax numbers look better, because of accumulated losses).
the historical average is closer to 7%. sustained 12% would be excellent growth for any mature firm
Data centers are real estate. One of the big players in carrier neutral data centers even calls themselves Digitial Realty.
The contents of the DC is not real estate. But neither is the an office or a house or a warehouse.
AND - they're operating well past their estimated service lifetime.
Not only that, but they're typically amortized over 5 years, where the actual lifespan usually falls far shorter (1-3 years), adding to the artificial subsidy conditions we see today. So they're gaming the lenders into deferring interest payments as much as possible today so that new competitors don't have the same cheap financing advantage.[0]
0: https://blog.citp.princeton.edu/2025/10/15/lifespan-of-ai-ch...
It's the sort of behaviour that really does end up with people going to prison.
There’s nothing fraudulent at all here just people using terms they really aren’t comfortable with.
That's the problem. That's the risk that few (if any) hyperscalers want to take.
The GPUs are far from worthless after 5 years. E.g. the A100 80GB PCIe version cost around $15k when it was introduced in 2021 and now sells for $10k used.
Things might be slightly worse for the data center servers, but I am sure they will find find buyers.
Which will not be any time soon according to SK Hynix CEO:
> We still forecast that customer demand will remain higher than our supply capacity even beyond 2030
https://www.reuters.com/world/asia-pacific/sk-hynix-ceo-sees...
Independent estimates sort of show around 2027-28 from what I remember.
I remember reading some article which said that RAM prices are already going down from its peak slowly (IIRC I can be wrong, I usually am but 3-5% month from its absolute peak) but the current RAM prices are still astronomical given past rates but the RAM prices will slow down hopefully sooner rather than later.
https://www.dpeaflcio.org/factsheets/the-professional-and-te...
In 4 years it better be 10x more important to have than a cell phone is today, or 10x more important than having internet/monitor/pc/printer is for an office worker today.
It's super-intelligence or bust.
Based on what? No AI company has ever made a cent in profit (exept for Nvidia lmao).
I hope you remember furniture.com, pets.com, webvan.com, kozmo.com and many others.
Amazon.com (during 1998) in some sense is the exception, not the norm from the bloodbath in stock markets during the dot com bubble.
(I highly recommend the book How the internet happened for more knowledge about things before, during and after the dot com bubble.)