1. First of all, the authors successfully trained a model with 173 BILLION PARAMETERS. The previous largest model in the literature, Google’s T5, had "only" 11 billion. With Float32 representations, GPT-3-173B's weights alone occupy ~700GB of memory (173 billion params × 4 bytes/param). A figure in the 100's of billions is still 3 orders of magnitude smaller than the 100’s of trillions of synapses in the human brain [a], but consider this: Models with trillions of weights are suddenly looking... achievable.
2. The model achieves competitive results on many NLP tasks and benchmarks WITHOUT FINETUNING. Let me repeat that: there is no finetuning. There is only unsupervised (i.e., autoregressive) pretraining. For each downstream NLP task or benchmark, the pretrained model is given text instructions, and possibly sample text with questions and answers. The NLP tasks on which the model was tested include translation, question-answering, cloze tasks, unscrambling words, using novel words in sentences, and performing 3-digit arithmetic.
3. The model is tested only in a ZERO-SHOT or FEW-SHOT setting. In other words, for each NLP task, the pretrained model is given text instructions with zero examples, or text instructions with a small number of examples (typically 10 to 100). As with human beings, GPT-3-173B doesn't need lots of examples to perform competitively in novel NLP tasks.
4. The results reported by this paper on all NLP tasks and benchmarks should be seen as a BASELINE. These results likely could be meaningfully improved with conventional finetuning.
5. The model’s text generation FOOLS HUMAN BEINGS, without having to cherry-pick examples.
--
[a] https://www.google.com/search?q=number+of+synapses+in+human+...