> The paper opens by saying "Large language models (with more than 1 billion parameters) perform well on a range of natural language processing (NLP) tasks in zero- and few-shot settings, without requiring task-specific supervision," and cite Radford et al. (2019), Brown et al. (2020), and Patwary et al., (2021) for this sentence. But the first paper doesn't claim zero-shot performance and the other two sources are about models orders of magnitude larger! I don't see any evidence in this paper to support the idea that there are people going around claiming that 1B+ parameter models have impressive zero-shot performance.
Maybe not zero-shot performance but they definitely claim few-shot performance which is what your quoted sentence is saying:
> Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. [1]
1: Language Models are Few-Shot Learners - https://arxiv.org/abs/2005.14165