> is the compute the main limiting factor today ?
I think it depends on the framing of the question, especially how you define compute.
- No, compute is not the limiting factor. The limiting factor is poorly optimized software (there's a joke: "10 years of hardware advancements have been entirely undone by 10 years of software advancements")
- No, compute is not the limiting factor. The limiting factor is that the electron is too big and the speed of light is too slow.
If we're talking about ML, then no, compute is not the limiting factor.
At least if we're define compute as the number of FLOPs we can process and not in terms of algorithm or resultant abilities. Though I'll admit that I'm an outlier in this respect[0]. But I think it is worth recognizing that we now have exaflop machines, that use tens of megawatts of energy, and they pale in comparison to what a 3 lb piece of fat and meat that only uses a handful of watts. In fact, our exascale computers aren't even seemingly sufficient to simulate far smaller and far less intelligent creatures. Certainly scale is a factor (we do see this pattern in apes too), but clearly there is more. And I think it should be obvious that scale isn't all you need, since we're the only ones. If it was that simple, we should see it more often. And if scale is indeed all we need, well we neither do we have an idea of how much scale that actually is nor does it mean that this is the best path forward as that scale may be ludicrously large. But what we do know, is that incredible feats can be done with what would constitute a rounding error to current scales (let alone future). I think we just want to believe this is the path forward because if it is, then there is a clear direction. But if it isn't, then we have to admit that we're still lost. But I think the problem is that we think that there's a problem in being lost. Or that we think that admitting we're lost somehow undermines or rejects the progress that we have made. But research is all about exploring the unknown. If you aren't at least a little lost, well then you're not exploring, you're reading a map. But the irony in this is that "scale is all you need" denies a lot of significant advancements we've made. Many smaller models perform far better that previously, and this is not due to knowledge transfer from larger models. Just look at any leaderboard, they aren't size is not the determining factor.
So I'd argue that if you want to advance AI, you should focus on smaller models. After all, smaller models are far easier to scale than larger models. They're also far easier to analyze and interpret, which is what gives us more information on how to lighten the way forward. But also don't expect a smaller model that is more successful to immediately be better than larger models. I far too often see a mistake even by reviewers/experts, where a method is dismissed because it was developed by some poor grad student with limited compute and did not unilaterally defeat the big models. Of course that doesn't mean the proposed methods are better, but that's orthogonal to what I'm arguing.
[0] Obviously I'm not alone. Yann is a clear believer and it's why he's looking at JEPA models (I don't think this will be enough but I think it is better). And Collet became more well known (at least outside the ML research community) and is a clear dissenter.