More than that, I think people overestimate how much AI will progress as you throw more compute at it. It’s the “9 women can’t deliver a baby in a month” equivalent of AI. Additional compute won’t magically give you AGI.
It's definitely too early to declare that more compute won't make a difference.
They were allegedly massive but the cost and returns were not worth it.
Of course, param count and context length are also important because they increase the model's overall fidelity, but a base model without SFT, RHLF etc is effectively useless.
Scale was really the unlock; the new pre and post training techniques and architectures are very cool and useful but they definitely aren't the differentiators when comparing to the previous era of NLP.
The fact that their advancement suggested that pouring more compute would continue working was also especially attractive to investors: it made a massive R&D budget feel like less of a risk.