That's why you capture multiple of them and verify your data statistically?
The problem is that what works on small LLMs does not necessarily scale to larger ones. See page 35 of [1] for example. A researcher only using the models of a few years ago (where the open models had <1B parameters) could come to a completely incorrect conclusion: that language models are incapable of generalising facts learned in one language to another.
The Twitter user doesn't even reference a single specific paper, kind of doing some hand wavy broad generalizations of his worst antagonists. So who really knows what he's talking about? I can't say.
If he means papers like the ones in this search - https://arxiv.org/search/?query=step+by+step+gpt4&searchtype... - they're all kind of interesting, especially https://arxiv.org/abs/2308.06834 which is the kind of "new prompt" class he's directly attacking. It is interesting because it was written by some doctors, and it's about medicine, so it has some interdisciplinary stuff that's more interesting than the computer science stuff. So I don't even agree with the premise of what the Twitter complainer is maybe complaining about, because he doesn't name a specific paper.
Anyway, to your original point, if we're comparing the research I linked and astronomy... well, they're completely different, it is totally intellectually dishonest to compare the two. Like tell me how I use astronomy research later in product development or whatever? Maybe in building telescopes? How does observing the supernova suggest new telescopes to build in the future, without suggesting that indeed, I will be reproducing the results, because I am building a new telescope to observe another such supernova? Astronomy cares very deeply about reproducibility, a different kind of reproducibility than these papers, but maybe more the same in interesting ways than the non-difference you're talking about. I'm not an astronomer, but if you want to play the insight porn game, I'd give these people a benefit of the doubt.