Agree that the AI bubble should pop though and the earlier, the better.
Agree that the AI bubble should pop though and the earlier, the better.
Even if they are heavily government subsidized for energy and hardware, I don see how the cost of training in the US would be more than double.
Not saying they are lying, but there incentives.
Everyone’s already begun trying this recipe in-house. Either it works with much less compute, or it doesn’t.
For instance, HKUST just did an experiment where small weak base models trained with DeepSeek’s method beat stronger small base models being trained with much more costly RL methods. Already this seems like it is enough to upend the low end models niche market, things like haiku and 4o-mini.
Be really skeptical why the people who should be making tons of money by realizing actually it was all a mirage and that they can now get the real stuff for even cheaper, would spend so much effort shouting about this, in order to undercut their own profitability..
tl;dr all numbers check up and the winnings come from the model architecture innovations they made.
Is it possible that they based their model architecture on the llama model architecture? Rather than just fine-tuned already training llama weights? In that case, they'd still have to do "bottoms up" training.