I think examples would be when identifying necessary and sufficient causes for a given effect.
A trivial case would be figuring out how fire works. You have have heat, fuel, and oxygen. Take any of those away, and there is no fire.
This can be demonstrated by direct experiment. But it’s unfortunately a pretty blunt tool. Because the real (and non-trivial) case is that this isn’t strictly true.
What you need is not specifically oxygen, but rather anything that can function as an oxidizing agent.
Still demonstrable by direct experiment, but more difficult. And how you want to understand fire is directly related to what you want to do with it: create it? Or make it go away?
If you want to make fire, it might be enough to understand the basic three element. You stumble onto the fact that magnesium will burn, and that’s good enough that you don’t look any further.
If you’re interested in putting fires out, it’s also “easy” to show that sumberging a burning object in water puts out the fire by direct experiment. Until you try to put out a magnesium fire like that. Similarly a phone battery that’s burning up will get worse if you dunk it in water.
Then on top of that, there are things that act like fire but are fundamentally different. Nuclear reaction as opposed to chemical ones.
Now, of course, we have a pretty good idea of how all of this works at a molecular or atomic level, but direct experimentation gets us a certain way down the path of investigating what’s really going on, even though it’s pretty imperfect.
But pretend that you have no knowledge of chemistry or physics, and you goal is to understand what causes fire. What you do know is some math and some statistics. How do you design a statistical approach to figuring out what makes fire? Or what makes fire go away?
You know that fire is related to heat because it’s hot. But you don’t know how. We also know that fire is related to wood because that our most frequent experience with fire. We think fire causes heat because everything that’s on fire is hot, and only some things not on fire are hot.
So there’s our null hypothesis: fire causes heat. The obvious alternative hypothesis is that heat causes fire.
And because we’re somewhat primitive in this example, we are really focusing on wood.
So we sample as many different types of burning wood as we can find. We go through the effort of finding samples of 1800 different types of tree on fire. (Doesn’t matter how the fire started, it just happened).
We find in every case that the burning wood was associated with heat of a certain temperature. So we accept the null hypothesis and reject the alternative.
Obviously, there’s a shit-ton of things that are obviously wrong with that statistical approach. And it probably seems comically bad to everyone. But the fact is that once you wade through the jargon specific to a given field, that’s really what you’re dealing with.
Null and alternative hypotheses that are causally confused, accepting rather than failing to reject . . . The whole thing. It happens mostly in fields where you cannot discover causal mechanisms through greater understanding of the platform. Social sciences, Econ, neuroscience, Climatology, etc.
And in the bad case above, note that the p-value would make no difference. The design itself and the misinterpretation is to blame. Not the p-value itself.
Yes, I think p-values and hacking them is a huge problem. But not because people have a habit of creatively interpreting data to get published. It’s because a focus on p-values encourages bad bad experimental designs, and reviewers don’t pay enough attention to that.