99% of the cost was in input tokens, I only used like 100k ish output tokens. It was a one shot task asking the agent to implement proxy injection to Guice. It did a pretty amazing job.
If you were to use hosted LLMs for a lot of agentic coding, a maxed out M5 Ultra Mac Studio would pay for itself in under a year.
Considering that I hit the 1M compaction multiple times per day with codex, it would definitely cost at least $5-8/day to use deepseek how I normally use codex.
That means all these data centers are being heavily utilized by actual end user inference demand. Well, some is research on new models, but a lot is actual end user demand. No one has given an explanation of why peoples usage would decline.
On top of that, margin on inference appears to be decent. It's model training that's a serious financial burden.
And maybe that's where there will be a slowdown, maybe the market doesn't justify spending as much on R&D as it does, but the end demand for inference is there.
Does that justify these stock prices? That's a different question. But the housing boom left behind endless rows of empty homes because demand disappeared. The 'dot com' boom left behind thousands of miles of dark fiber that'd been built out well ahead of demand for bandwidth. I can see the stock market having a giant sell off, but I don't see data centers sitting idle in that same fashion.
Here's what you have to believe:
- AI demand is at least several times larger than what can currently be satisfied, or will grow. (This one I can buy, but...)
- AI chips (GPUs, TPUs, compute-in-memory, whatever else is being studied) will not get significantly more efficient than they are now. It will not be possible in, say, 5-10 years, to do 2X or 4X or 10X more AI requests per rack than is possible now. I think this one's the single most likely thing to be false, since all computing history contradicts it.
- Edge devices (PCs, laptops, specialized but smaller scale AI compute nodes) will never be powerful enough to run frontier models at a reasonable price that's appealing for professionals, enthusiasts, or businesses, and there will never be a market for this. None of the demand will be served on-device or near-edge. AI must all go in giant data centers.
- AI models will not become significantly more efficient than they are now. There are no large gains on the table from better model architectures, better training, more efficient quantizations, better harnesses, etc.
If all those things are true, than the current planned like 4X-10X increase in data center capacity makes sense. If even one or two of them are not true, then the planned data center build-outs start looking excessive. If all four are not true, it's a total bubble that will crash and burn. Answer is probably somewhere between, but how far toward bubble? That's why I picked a number like "only 20% ever gets built." It might be as high as 50%. It ain't gonna be 100%. The planned built-out is batty.
Oh I forgot two more...
- Data center capacity currently serving non-AI work loads does not shrink through either reduced demand, more efficient software, or (most likely) faster chips and denser RAM. If that happens, more pre-existing DC space can serve AI work loads.
- Orbital solar powered compute nodes never happen. If this happens (free power! much less political opposition!) then terrestrial data centers have significant competition.