It's the Effect Size, Stupid: What effect size is and why it is important (2002)
leeds.ac.uk
leeds.ac.uk
Of course, effect size isn't enough either. You still need a good understanding of the range of outcomes, if there are significant fraction of participants who had poor/negative outcomes, etc. This can be hidden a bit inside a single metric like effect size.
If you have just one white lab rat (N=1) and it weighs 50 kg after the treatment, you can just skip all statistical analysis an publish the result. The importance of the discovery is self evident due to the effect size alone.
The discovery that H. pylori causes peptic ulcers was first submitted to Lancet as a two papers, 25 patients study and 100-patient study. Lancet was slow to publish and there was resistance from the medical establishment. It was hard to find reviewers who would agree on the importance of the paper. Barry Marshall decided to do experiment with himself and it was very convincing. Marshall and Warren received a Nobel price for their discovery.
We would encounter this a lot in performance engineering - sometimes you could isolate that a new optimization was statistically significantly better, yet only worth 0.1% in terms of performance improvement (possibly in exchange for a whole bunch of new code).
https://docs.google.com/spreadsheets/d/1eEBFGRO1UgA6OYoUF9Xg...
Take effect size, for example. Suppose I run my test before and after the latest commit. The software now has 2% fewer TPS.
Is that meaningful? It depends. Should I rely on a single run? Almost certainly not.
Suppose performance improves by 203%. Is that meaningful? Probably. Should I rely on a single run? Almost certainly not.
Then there's the obnoxious problems of vast amounts of uncontrollable variables. Folks run buckets of testing on cloud platforms without knowing if they're landing on a physical machine they used last time and so their bits are disk-warm, or that there's a network glitch in central1-a but not central1-b but they don't know which one they're using, or they run tests at different times of day and don't realise they will get different competition for cloud resources due to diurnal demand ...
tl;dr run more tests and get someone to help you. If you run ab2 three times and publish a breathless blog post, be that on your soul.