It would seem that this can be easily analysed scientifically.
To give a simple example: if, hypothetically, someone thought that GPT-3 is good at basic arithmetic (1 plus 1, 1000 times 3 etc.), they can provide a template for how to ask GPT-3 questions about arithmetic. Anyone can then verify that this template results in accurate answers, by asking randomly sampled questions using that template.
This verification method could be applied to pretty much any problem. Has anyone done anything like that?