Perhaps we have to distinguish between GPT-3-the-model and GPT-3 paper. IMHO GPT-3 as a model is straighforward engineering, putting a lot of resources in an oversized GPT-2; and while there's significant novelty in the "Language Models are Few-Shot Learners" paper about how exactly you apply these models,
that is orthogonal to GPT-3-the-model, the scientific content of that paper applies to any other powerful language models and isn't intimately tied to specifics of GPT-3.
In essence, I feel that the same people introduced two quite separate things - a completely new paradigm on how to obtain few-shot learning from a language model in a way that competes with supervised learning of the same tasks; and the GPT-3 large model which is used as "supplementary material" to illustrate that new paradigm bit is also usable and used with the old paradigms, and by itself isn't a breakthrough. And IMHO when the public talks about GPT-3, they do really mean GPT-3-the-model and not the particular few-shot learning approach.