All major cloud providers have high profit margins in the range of 30-40%.
How much does it cost to train a cutting edge LLM? Those costs need to be factored into the margin from inferencing.
Buying hard drives and slotting them in also has capex associated with it, but far less in total, I'd guess.
How much does it cost to train a cutting edge LLM? Those costs need to be factored into the margin from inferencing.
They don't, though! I can buy hardware off of the shelf, host open source models on it, and then charge for inference:Where is the return on the model development costs if anybody can host a roughly equivalent model for the same price and completely bypass the model development cost?
Your point is inline with the entire bear thesis on these companies.
For any use cases which are analytical/backend oriented, and don't scale 1:1 with number of users (of which there are a lot), you can already run a close to cutting edge model on a few thousand dollars of hardware. I do this at home already
DeepMind is actively using Google’s LLMs on groundbreaking research. Anthropic is focused on security for businesses.
For consumers it’s still a better deal for a subscription than to invest a few grand in a personal LLM machine. There will be a time in the future where diminishing returns shortens this gap significantly, but I’m sure top LLM researchers are planning for this and will do whatever they can to keep their firm alive beyond the cost of scaling.
I am not suggesting these companies can't pivot or monetize elsewhere, but the return on developing a marginally better model in-house does not really justify the cost at this stage.
But to your point, developing research, drugs, security audits or any kind of services are all monetization of the application of the model, not the monetization of the development of new models.
Put more simply, say you develop the best LLM in the world, that's 15% better than peers on release at the cost of $5B. What is that same model/asset worth 1 year later when it performs at 85% of the latest LLM?
Already any 2023 and perhaps even 2024 vintage model is dead in the water and close to 0 value.
What is a best in class model built in 2025 going to be worth in 2026?
The asset is effectively 100% depreciated within a single year.
(Though I'm open to the idea that the results from past training runs can be reused for future models. This would certainly change the math)
It doesn't even seem like these companies are in a battle of attrition to not be the first to go bankrupt. Watching this would be a lot more exciting if that was the case! I think if there was less competition between LLMs developers could slow down, maybe.
Looking at the prices of inference of open-source models, I would bet proprietary models are making a nice margin on API fees, but there is no way OpenAI will make their investors whole because they make a few dollars of revenue for a million tokens. I am terrified of the world we will live in if OpenAI will be able to reverse their balance sheet. I think there's no where else that investors want to put their money.
In the end it comes all down to the value provided as you see in the storage example.
If inference has significant profitability and you're the only game in town, you could do really well.
But without regulation, as a commodity, the margin on inference approaches zero.
None of this even speaks to recouping the R&D costs it takes to stay competitive. If they're not able to pull up the ladder, these frontier model companies could have a really bad time.
Of course that means it's unprofitable in practice/GAAP terms.
You'd have to have a pretty big margin on inference to make up for the model development costs alone.
A 30% margin on inference for a GPU that will last ~7 years will not cut it
I'm not sure about how the regulation of things would work, but prompt injections and whatever other attacks we haven't seen yet where agents can be hijacked and made to do things sounds pretty scary.
It's a race towards AGI at this point. Not sure if that can be achieved as language != consciousness IMO
Who is "we", and what are the actual capabilities of the self-hosted models? Do they do the things that people want/are willing to pay money for? Can they integrate with my documents in O365/Google Drive or my calendar/email in hosted platforms? Can most users without a CS degree and a decade of Linux experience actually get them installed or interact with them? Are they integratable with the tools they use?
Statistically close to "everyone" cannot run great models locally. GPUs are expensive and niche, especially with large amounts of VRAM.
I'm not saying the options are favorable for everybody, I'm saying the options are there if it becomes locked in to 1-3 companies.
However it is arguable that thought is relatable with conscienceness. I’m aware non-linguistic thought exists and is vital to any definition of conscienceness, but LLMs technically dont think in words, they think in tokens, so I could imagine this getting closer.
No one accuses the Dewey decimal system of thinking.
Sorry.
It is a complex supply chain but each section of the chain is held by only a few companies. Hopefully this is enough competition to accelerate the development of computational technologies that can run and train these LLMs at home. I give it a decade or more.