It's also worth keeping in mind that depending on benchmark, these values change (and can shrink quite a bit)
And it's also worth keeping in mind that the drastic drop in training cost(if reproducible) will mean that training is suddenly affordable for a much larger number of organizations.
I'm not sure the impact on GPU demand will be as big as people assume.
In the long run, yes, they will be cheaper due to more competition and better tech. But next month? It will be more expensive.
It'll be much harder to convince people to buy the latest and greatest with this out there.
The usage of existing but cheaper nvidia chips to make models of similar quality is the main takeaway.
So why not buy a more expensive Nvidia chip to run a better model?The DeepSeek R1 model people are freaking out about, runs better with more compute because it's a chain of thoughts model.
The analogous costs would be what OpenAI spent to go from GPT 4 to GPT 4o (i.e., to develop the reasoning model from the most up-to-date LLM model). $5 million is still less than what OpenAI spent but it's not a magnitude lower. (OpenAI spent up to $100 million on GPT4 but a fraction of that to get GPT 4o. Will update comment if I can find numbers for 4o before edit window closes)