> No, if you knew the null model was flawed to begin with nothing has been learned.
We don't know that the null is flawed to begin with. Not in a way that matters to the analysis anyway, otherwise you've picked the wrong null. No one runs analyses to discover the pion for the millionth time.
For BSM work, the way it generally works is this (ignoring limit setting):
- Pick a signature predicted by one or more BSM models (or even none), that either isn't predicted by the SM or is exceedingly rare in the SM. For example, you might predict a new particle that decays in some predetermined way.
- Define a signal region in some combination of variables where you expect to find a signal corresponding to this signature. This region is blinded, and you don't look at it until after fixing your entire procedure.
- Estimate the background in this signal region with a combination of Monte Carlo simulations and extrapolation from outside this region (data-driven backgrounds). This obviously has both statistical and systematic uncertainties, and a lot of the hard work is in getting this right. For instance, you might use independent Monte Carlo generators, or define control regions to check your estimation in.
- "Open the box" and look at the signal region data. If you see an excess over the estimated background, calculate the significance. Even here, if you were looking for a new particle, for instance, and the excess isn't a localized bump, this would be an indication that the background estimation may be flawed.
- If the significance is over 5 sigma, you have a discovery.
As you can see, you aren't using a null hypothesis that you already know is flawed. Your null is specifically "no signal exists". A positive deviation that isn't a statistical fluctuation is by definition a discovery of a signal.
Things like malfunctioning equipment go into the uncertainties if they occur in a way that the researchers have considered (which often means uncertainties are set conservatively if the equipment behaviour is poorly understood). If they occur in a way that no one considered, there's no statistical trickery that's every going to compensate for that. If we just went straight to setting limits, we would still see a deviation from what's expected there if equipment malfunctioned in a manner that faked a signal.
> In summary: If there is anything wrong with the null/background model at all, the p-value will converge on zero, it is just a matter of collecting enough data.
In summary: Assuming the background estimation is correct, the p-value will only converge to zero if a true signal exists. If the background estimation is wrong, then the background estimation is wrong and this is a problem with the background estimation, not hypothesis testing.
I'm by no means claiming the system is perfect, but it isn't systematically flawed the way you seem to be claiming.