As far as I can tell, it is a combination of smart algorithms, good engineering, and the hardware to make it happen. And scientists that had the right hunch for which direction to push in.
Well, GPT-3 isn’t a classifier and it isn’t using labeled data.
As an outsider it definitely appears that GPT-3 is an engineering advancement, as opposed to a scientific breakthrough. The difference is important because we need a non linear breakthrough.
GPT-3 is a bigger GPT-2. As far as we know, there is no more magic. But I think it’s a near certainty that larger models will not get us to AGI alone.
In essence, I feel that the same people introduced two quite separate things - a completely new paradigm on how to obtain few-shot learning from a language model in a way that competes with supervised learning of the same tasks; and the GPT-3 large model which is used as "supplementary material" to illustrate that new paradigm bit is also usable and used with the old paradigms, and by itself isn't a breakthrough. And IMHO when the public talks about GPT-3, they do really mean GPT-3-the-model and not the particular few-shot learning approach.
Quality is increasing with parameters. Even now, interfacing with codex leads to unique and clever solutions to the problems I present it.
So is the answer here both?
Some scientist at OpenAI had the hypothesis "what if our algorithm is correct, but all we need is to scale it 3 magnitudes larger" and made it happen. They figured out how to scale all the dimensions. How fast should x scale if I scale y? (That is very tricky, as modern machine learning is basically alchemy)
And then they actually scaled it. That took a ton of engineering and hardware, for what was essentially still following a hunch.
And then they actually noticed how good it was, and did a ton of tests with the surprising results we now all know.
I'm not sure many would agree that the desire to scale was simply a "hunch".
GPT-3 has 175B parameters. The previous largest model was Microsoft's Turing-NLG which had 17B. GPT-2 had 1.5B.