I don’t see this as good evidence that model parameter increases have reached a limit.
I see it as evidence that compute costs are very high.
I don’t see this as good evidence that model parameter increases have reached a limit.
I see it as evidence that compute costs are very high.
GPT-3 cost 4.6 million dollars to train. According to a cursory Google search, GPT-4 cost 100 million dollars to train, about 50x as much. Is a 0.3x improvement worth 50x the cost?
That doesn't even touch on the costs of inference. Millions of queries a day requiring heavy duty compute infrastructure. How much capital are we going to waste on the next marginal improvement? Perhaps Microsoft will "only" need another 30% increase in their carbon emissions this year to sustain the growth.
There are many things which earlier models maybe somewhat did, but only large models do reliably to the point of being usable for more than tech demos
I find gpt-4 and gpt-3 a world apart. Perhaps it's just that it crossed a tipping point that made it useful for _me_ and the next 10x investment in training wouldn't achieve the same effect for me (but perhaps for somebody else?)
He also says
>Paradoxically, smaller models require more training to reach the same level of performance.
It's not a paradox at all. If less training to reach the same level of performance was true, that would be a paradox whereby they'd be trained for under a nanosecond to achieve optimal performance/size payoff.