The number of trials is one issue, but another is selection bias. If GPT-4 got a worse result, would he have tweeted it? And if he did tweet it, would people like it/upvote it to the front of HN? Less likely given hype goes more viral than doubt. So they need to publicize the methodology of the experiment ahead of time, then run the experiment after the disclosure. Ideally, the intention to run the experiment gets to the front page of HN prior to it being run, otherwise we've still got some selection bias at play.