From the original paper:
"Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model."
Of course, it makes sense that people critique GPT-3 in terms of its actual progress towards human-like intelligence, since every news publication writes about it like it's Skynet and OpenAI's stated goal is AGI. I do think, however, that we miss all the parts of GPT-3 that are exciting and innovative when we view it through the binary of "Is it really understanding as a human does?" (Not that you were reducing it to that, just speaking about the conversation around GPT-3 more broadly.)