https://openai.com/index/scaling-laws-for-neural-language-mo...
And a return to more normal world where progress comes from refinement of algorithms and approach rather than brute force compute.
Deepseek is just the poster child.
https://openai.com/index/scaling-laws-for-neural-language-mo...
And a return to more normal world where progress comes from refinement of algorithms and approach rather than brute force compute.
Deepseek is just the poster child.
Thermodynamic neural networks may also basically turn everything on its ear, especially if we figure out how to scale them like NAND flash.
If anything, I would estimate that this is a space-race type effort to “win” the AI “wars”. In the short term, it might work. In the long term, it’s probably going to result in a massive glut in accelerated data center capacity.
The trend of technology is towards converging with or doing better than natural processes, not doing it 100000x less efficiently. I don’t think AI will be an exception.
If we look at what is -theoretically- possible using thermodynamic wells, with current model architectures, for instance, we could (theoretically) make a network that applies 1t parameters in something like 1cm2. It would use about 20watts, back of the napkin, and be able to generate a few thousand T/S.
Operational thermodynamic wells have already been demonstrated en silica. There are scaling challenges, cooling requirements, etc but AFAIK no theoretical roadblocks to scaling.
Obviously, the theoretical doesn’t translate to results, but it does correlate strongly with the trend.
So the real question is, what can we build that can only be done if there are hundreds of millions of NVIDIA GPUs sitting around idle in a few years? Or alternatively, if those systems are depreciated and available on secondary markets?
What does that look like?
Bigger models are more capable, but smaller models can be iterated on faster. For a couple months now most of the impressive achievements have been in increasingly smaller and cheaper models. Deepseek just has the perfect storm of impressive results, accessibility and international rivalry that made it go viral
It's actually the beginning of test time scaling. R1 has shown that a very simple reinforcement learning scheme can be used to teach the model how to think in a chain-of-though as an emergent property.
No addition pretraining data needed! Only more compute.