Better in this case means some combination of "less errors for the same size" and/or "bigger and smarter". Fundamentally, they're still the same thing, just more and better.
Unfortunately, the scaling is (roughly) logarithmic. So for every 10x increase in scale you get a +1 better model. Scaling up 1,000x gets you just a +3 improvement, and so on.
This is useful to eke out every last drop of quality per gigabyte of model file size. It also keeps the models up to date with current events.
Obviously this scaling becomes too inefficient at infinite scale not just because of training costs that’ll never be recouped but also increasing inference cost with larger models.
Some fundamentally new architectures will need to be developed to take much better advantage of increased computer power.
I suspect the major players are investing in hardware now in the hope that some revolutionary new algorithm is invented soon and they’ll be ready for it.
It’s… a bit of a gamble!
Bleed investors dry before the next fad pops up
A G650 to fly to your 85m yacht in the med doesnt come cheap.