Further down there's a plot titled "Aggregate Performance Across Benchmarks" where we can see that performance is on average about 50%. I don't know what the baseline for this plot should be (what is the expected average if the tasks are solved by a random classifier?) but comparing this plot with the plot at the top of the article, it doesn't really look like there's a huge improvement in accuracy with a huge increase in the number of parameters. In fact, it's quite the contrary: there's a small increase and a very smooth, almost linear curve. So that's an exponential increase in the use of resources for an almost linear increase in performance? That's not that impressive.
So it appears that the big thing about GPT-3 is that it's big.
It should also be noted that the public interest about GPT-3 is mostly focused on its ability to generate text, for which there is no good metric. So basically, GPT-3 is big, but it's not that good in tasks for which there are formal benchmarks (such as they are, because Natural Language Understanding benchmarks are often very poorly made and don't really measure what they say they measure) and we can't really tell how good it is in the one task that interests most people the most.