AI models that cost $1B to train are underway, $100B models coming
tomshardware.com
tomshardware.com
$100m is manageable, if you've got 100m paying subscribers or companies using your API for a year you can recoup the costs, but there aren't many companies with 100m users to monetise for it. $1B feels like it's pushing it, only a few companies in the world can monetise, and realistically it's about lasting through the next round to be able to continue competing, not about making the money back.
$100B though, that's a whole different game again. That's like asking for the biggest private investment ever made, for capex that depreciates at $50B a year. You'd have to be stupid to do it. The public markets wouldn't take it.
Investing that much in hardware that depreciates over 5+ years and is theoretically still usable at the end, maybe, but even then the biggest companies in the world are still spending an order of magnitude less per year, so the numbers end up working out very differently. Plus that's companies with 1B users ready to monetise.
For AGI, the bet is that someone will build an AI capable enough to automate AI development. Once we get there it will pay for itself. The question is what the cost-speed tradeoff to get there looks like.
This seems less like doing business as usual and more like betting big to be part of something really transformative.
it may never happen, especially with this current approach
at which point you've burnt hundreds of billions of dollars, emitted millions of tonnes of CO2 and all you've got out of it is a marginally better array of doubles
Isn’t that exactly what’s happening?
A $300k 8x H100 pod with 5kW power supply burns at most $6k per year in power at $0.15/kWh. The majority of the money is going to capital equipment for the first time in the software industry in decades.
These top of the line chips last for much longer in the depreciation game. The A100 was released in 2020 but cloud providers still have trouble meeting demand and charge a premium for them.
NVIDIA claim 10.2kW for a DGX H100 pod. https://docs.nvidia.com/dgx/dgxh100-user-guide/introduction-....
Your point still stands though where power is a fraction of the cost.
The bigger issue is power + cooling and how many units are needed to train the better models.
If a 100 Billion dollar training run produces the highest quality model in the land across all metrics and capabilities. That will be the model thats used, at most there would be 1-2 other firsm willing to spend 100 Billion to chase the market.
If I can come out with a model a year later, and it can provide 95% of the performance while costing 10% as much to run, I think I would end up stealing a lot of customers before they had a chance to break even.
Take Llama3-8B for example, this is an 8 billion parameter model from 2024 that performs about as well the the original ChatGPT, a 175 billion parameter model from 2022. It only took 2 years before a model that can run on a desktop could compete with a model that required a data center.
If you're MSFT - you don't care who wins as long as you have cost competitive rights to embed the AI in all of your products - earlier than others.
This is the Anthropic CEO talking up his company's capital needs to the Norwegian Sovereign Wealth Fund ( Norges Bank Investment Management ) and trying to justify some absurd 100bn valuation.
While I don't disagree 100%, my question to you is:
who/what says this is the case/why? GPT-3.5 was released/made popular "to the masses" not too long ago. Where do you feel the pressure for a quantum leap "quickly" is coming from?
> That’s true for AI, but it is not the right way to think about AGI. For AGI, the bet is that someone will build an AI capable enough to automate AI development. Once we get there it will pay for itself. The question is what the cost-speed tradeoff to get there looks like.
I don't think people are treating AI as a typical investment. They are valuing it as a potential replacement for like 95% of human workers. Once the plateau becomes obvious to even the biggest fanatics, people are going to realize all this money has been used to create really good chatbots that just make shit up 25% of the time. The whole sales pitch for the last 1-2 years has been that AGI is "just around the corner" and we can get there via the magic of exponential growth.
It's entirely possible we could see an AGI develop within a specific medium, e.g. text-only to start.
But also, don't you think those two problems are convergent? If we make an algorithm smart enough to drive a car anywhere more safely than a human would, it seems likely to me that that will either be AGI or right on the cusp of it.
I like to play a game whenever I meet a new group of programmers/lawyers/board gamers/pedants. I ask them to define a sandwich, then I point out all of the edge cases until we get to the point where I can either say their definition is bad because it e.g. doesn't include subs (short for submarine sandwiches), or I can say that poptarts are sandwiches.
The problem is that reality is blurry and clear definitions are fundamentally incapable of capturing the nuance there.
See also standardization of EV battery form factors -- a problem that, had it been tackled by government several years ago, would have avoided the chicken-and-egg problem that is impeding adoption now.
Doing automated driving right means restructuring our entire surface transportation system to support it. I doubt it can happen without a serious commitment from government... which is in no condition to commit to anything.
I am pretty sure that's less than half the sales pitch. Current LLMs are economically valuable, they are useful to all kinds of knowledge work, despite their flaws.
Or is it the total overall cost of buying TPUs / GPUs, developing infrastructure, constructing data centers, putting together quality data sets, doing R&D, paying salaries, etc. as well as training the model itself? I could see that overall investment into AI scaling into the tens of billions over the next few years.
Better in this case means some combination of "less errors for the same size" and/or "bigger and smarter". Fundamentally, they're still the same thing, just more and better.
Unfortunately, the scaling is (roughly) logarithmic. So for every 10x increase in scale you get a +1 better model. Scaling up 1,000x gets you just a +3 improvement, and so on.
This is useful to eke out every last drop of quality per gigabyte of model file size. It also keeps the models up to date with current events.
Obviously this scaling becomes too inefficient at infinite scale not just because of training costs that’ll never be recouped but also increasing inference cost with larger models.
Some fundamentally new architectures will need to be developed to take much better advantage of increased computer power.
I suspect the major players are investing in hardware now in the hope that some revolutionary new algorithm is invented soon and they’ll be ready for it.
It’s… a bit of a gamble!
Bleed investors dry before the next fad pops up
A G650 to fly to your 85m yacht in the med doesnt come cheap.
I wonder which timelines had this scenario…
they would be better off not bullshitting their investors.
investors with huge piles of cash should buy themselves a brain and stop funding bullshitters