For instance, the Higgs is part of the Standard Model but for the Higgs discovery the null hypothesis was that the Higgs does not exist.
For instance, the Higgs is part of the Standard Model but for the Higgs discovery the null hypothesis was that the Higgs does not exist.
>"Assuming the background estimation is correct, the p-value will only converge to zero if a true signal exists."
Aren't you assuming the background model is correct and a true signal exists?
>"Your null is specifically "no signal exists"."
No, it is never this. It is an entire model, one assumption of which is "no signal exists". The p-value doesn't care which assumption is wrong, it is only wishful thinking that leads to people focusing on that one.
>"We don't know that the null is flawed to begin with."
Sorry, I don't believe this. Give a concrete example. Do you believe the standard model is 100% correct? If not, then any background derived from it must be assumed to be flawed to begin with.
>"Things like malfunctioning equipment go into the uncertainties if they occur in a way that the researchers have considered (which often means uncertainties are set conservatively if the equipment behaviour is poorly understood). If they occur in a way that no one considered, there's no statistical trickery that's every going to compensate for that."
Correct, there is no statistical trickery to compensate for this. It is a scientific problem, not statistical.
Did you check out that Meehl paper I linked elsewhere?[1] If you take a prediction of a model and test that, messing up the experiment means you get results that diverge from your prediction. If you reverse the logic of science so that you test the "opposite" of your prediction, messing up the experiment yields results that seem to support your model. This is why the prediction of the theory needs to be set as the "hypothesis to be nullified".
[1] https://meehl.dl.umn.edu/sites/g/files/pua1696/f/074theoryte...
I know you're feeling very right and vindictive but damn if this isn't tiring to read. Make sure you know what you're right about first before assuming the other thing is wrong just because it matches a few keywords that you've read are bad. Because assuming good faith means that you allow for the possibility they might not be doing it wrong. You still get to make your point but it saves your face a lot.
I'd suggest, go outside, attack a windmill for a few hours, think about what you're actually trying to say, and if it's not "hey I read this article on lesswrong and it says everybody is doin it wrong", you'll probably find you can make your point much clearer.
You seem to have more than sufficient knowledge about statistics to think for yourself on this matter.
Quote this. Where? I am sure by "corrected" you mean "someone claimed something I believe to be wrong, so I continued to point out why." Amazingly, this looks like yet more strawman arguing...
1. If you're using a data driven background, you're background model isn't the Standard Model but that there isn't something special about your signal region.
2. If you are using the SM as your background, then yes you do believe it to be 100% true, since what you are looking for is evidence that there exists new physics not described by the SM.
In other words, the theory under test is the SM, and it can pass very easily. If it fails despite this, that's evidence that it's incomplete. If it passes, the limit setting stage tests whatever new physics we were looking for, and gives an upper bound on how sure we are that it does not exist.
Sure, it means you assume that during the signal detection everything was working exactly as it was during background data collection. There are going to be other implementation specific assumptions like what thresholds to use regarding environment noise triggers, accounting for sensor drift, etc. These data-driven models are approximations, I doubt the people actually coming up with them really believe they are 100% true. If you give a real life example I will point out exactly where there are issues.
> "If you are using the SM as your background, then yes you do believe it to be 100% true"
I don't know what this is supposed to mean. Of course you are assuming your background is 100% true. Do particle physicists actually believe this though? Not from my reading:
"So although the Standard Model accurately describes the phenomena within its domain, it is still incomplete." https://home.cern/about/physics/standard-model
As for the SM, we are looking for those places where it is incomplete, and to do this we must assume it is complete and look for strong evidence that we are wrong. To date, it doesn't seem like we see though, as far as collider physics is concerned.
Yes, I know. This isn't at all a positive thing as you seem to think...
I've been lately seeing that this practice is infecting particle physics, which is not a good sign. I do not think they are overwhelmed yet though. Here is how it has gone in every field so far (education research, psychology, medicine, etc):
Eventually the field will stop caring about all actual quantitative predictions (some barely funded crackpots on the edges will publish in ignored journals) and researchers will only test a "null hypothesis" that everyone knows is wrong (eg, in the physics case a world where the standard model is assumed true but has no higgs boson). Then upon rejecting the (known to be false) null hypothesis, they will conclude some theory only capable of vague/malleable predictions is correct.
The way it works is simple, the null model acts as a strawman. The confidence in rejecting the null model is somehow translated into support for the researcher's favorite theory. How does changing the significance level address this problem?
Not just any old favorite theory. For instance, an excess at a mass of 125 GeV doesn't provide any evidence for a new particle at 2 TeV. It does however provide evidence for a new particle at 125 GeV.
> Eventually the field will stop caring about all actual quantitative predictions (some barely funded crackpots on the edges will publish in ignored journals) and researchers will only test a "null hypothesis" that everyone knows is wrong (eg, in the physics case a world where the standard model is assumed true but has no higgs boson). Then upon rejecting the (known to be false) null hypothesis, they will conclude some theory only capable of vague/malleable predictions is correct.
If that were the case, we would have made a lot more discoveries by now.
In physics you still have quantitative predictions done in addition to checking the null hypothesis (which I why I said that was superfluous), if NHST adoption continues these will slowly be deprecated and replaced by only checking the null hypothesis. Seriously go read some old psychology before NHST destroyed it, they were developing quantitative theories of learning and everything [1,2]. Then go look at it now (statistically significant this and that with basically no progress for 60 years...).
Also, the significance level, constraints on p-hacking, etc are adjusted so that discoveries occur at the "right" rate. I'm not sure what triggered he move to 5 sigma in physics but it was probably something like "this seems too easy" so people started carefully double checking and getting conflicting results.
[1] https://www.tandfonline.com/doi/pdf/10.1080/00221309.1934.99...
Nevertheless, I suppose what you would suggest is just setting limits to start with and looking for a bump there instead. That's not a bad idea, and some (Gross, e.g.) have suggested setting limits even in discovery papers.
I don't see how this addresses any of the issues, it just makes it more expensive (ceteris paribus). As I already wrote:
"A more stringent significance level just means it costs more money or the data is much cheaper to collect. ...The way it works is simple, the null model acts as a strawman. The confidence in rejecting the null model is somehow translated into support for the researcher's favorite theory. How does changing the significance level address this problem?"
>"Nevertheless, I suppose what you would suggest is just setting limits to start with and looking for a bump there instead."
Exactly, and the more precise the limits, the more meaningful for the theory if its consistent with the data. At the extreme, theories that allow for any value at all (as I understand the case is for every possible observation under string theory) are not supported by the data no matter how it turns out.
Confidence in rejecting the null isn't "somehow translated into support for the researchers' favorite theory", it's translated into support for the existence of a signal, which makes sense since the null is the background-only hypothesis.
After we establish rudimentary support for a model by finding one or more such signals, we can then embark on the process of doing precision measurement to really test that model (as is being done for the SM).
If its only a statistical fluctuation, the p-value due to sequential sampling follows a markov process (random walk) between zero and one (there is no convergence). So starting from a significant result, the p-value would be just as likely to decrease and become "more significant" as increase and "lose significance". The significance cutoff is irrelevant to this behavior, although the random walk will spend less time in these extreme regions to begin with.
Here is an example of the random walk in R. It was my first attempt when going up to 100k samples, after n = 45 the p-value was ~5e-4 (~3.3 sigma), later at n = 36441 it went even more extreme in the other direction, corresponding to a p-value of ~1.9e-6 (~4.65 sigma). Here is the plot and code: https://image.ibb.co/hOHO08/p_markov_chain.png
set.seed(1234)
n = 1e5
a = rnorm(n)
b = rnorm(n)
p = sapply(10:length(a), function(i) t.test(a[1:i], b[1:i])$p.value)
plot(p, type = "l", panel.first = grid())
> min(p)
[1] 0.0005394751
> 1 -max(p)
[1] 1.908121e-06
Perhaps I'm wrong and if you run enough of these and subset to only look after a certain threshold has been breached (ie 5e-2 or 3e-7) there will be different behavior, I doubt it though. However, that is all unimportant since it is based on a false premise:>"_if_ what you're seeing is a statistical fluctuating"
You aren't you are seeing that, so everything that follows is irrelevant. The background model is imperfect (eg nobody believes in a standard model with no higgs boson, everyone knows it is wrong somehow). In this (realistic) case convergence to statistical significance is guaranteed with enough data. However, in the rare (non-existent?) case where the background model is thought to be perfect, it is a different story.
>"Confidence in rejecting the null isn't "somehow translated into support for the researchers' favorite theory", it's translated into support for the existence of a signal, which makes sense since the null is the background-only hypothesis."
This amounts to claiming that in most peoples minds (including the authors of the paper) the Higgs boson, gravitational waves, etc haven't been "detected with 5 sigma confidence"... How much would you want to bet on this? The null model could be rejected due to a loose cable (ftl neutrinos), but the p-value doesn't distinguish between these explanations. Perhaps there was "a loose cable" in the LHC experiments but nobody looked very hard since the deviation agreed with preconceived notions? The p-value doesn't help you here.
>"After we establish rudimentary support for a model by finding one or more such signals, we can then embark on the process of doing precision measurement to really test that model (as is being done for the SM)."
Why not just skip the first step? Instead collect data to check whatever models that exist. When none of them can explain it, modify or come up with new models until it fits. Then derive new predictions from those and collect new data to check them, repeat.
EDIT: I ran it again and got this: https://image.ibb.co/g9rGL8/p_markov_chain2.png
> min(p)
[1] 0.005265033
> 1-max(p)
[1] 1.453814e-05If you collect, say, 1000 samples of a normally distributed variable and find that the mean (for instance) is 2 sigma from the expected mean, then if you collect another 1000 samples and take the mean of all 2000 samples, you will probably find that the significance decreases, assuming the first result was a statistical fluctuation. This is just [regression to the mean](https://en.wikipedia.org/wiki/Regression_toward_the_mean).
What we don't do is dishonestly drip-feed data until we hit the significance we want, then stop. Data is added in fairly large chunks (generally an year's run at a time), and when a new dataset is added the analysis is developed with (at least the new) data blinded. The dataset is only unblinded when the analysis has been finalized.
You have found a deviation from your null model. That is it.
>"so it's right to claim a discovery."
No, if you knew the null model was flawed to begin with nothing has been learned. Also I wouldn't refer to something like "equipment was malfunctioning" as a discovery. Drawing any conclusions about the actual research hypothesis must be done outside the NHST framework, or else there is a fallacy at play (probably strawman).
>"What we don't do is dishonestly drip-feed data until we hit the significance we want, then stop. Data is added in fairly large chunks (generally an year's run at a time), and when a new dataset is added the analysis is developed with (at least the new) data blinded. The dataset is only unblinded when the analysis has been finalized."
Great, but even if you don't p-hack it still doesn't work unless it really makes sense to assume the null/background model is perfectly true. P-hacking is just more BS on top of an already BS procedure.
In summary: If there is anything wrong with the null/background model at all, the p-value will converge on zero, it is just a matter of collecting enough data. This will happen regardless of whether the theory of interest (research hypothesis) is accurate or not.
We don't know that the null is flawed to begin with. Not in a way that matters to the analysis anyway, otherwise you've picked the wrong null. No one runs analyses to discover the pion for the millionth time.
For BSM work, the way it generally works is this (ignoring limit setting):
- Pick a signature predicted by one or more BSM models (or even none), that either isn't predicted by the SM or is exceedingly rare in the SM. For example, you might predict a new particle that decays in some predetermined way.
- Define a signal region in some combination of variables where you expect to find a signal corresponding to this signature. This region is blinded, and you don't look at it until after fixing your entire procedure.
- Estimate the background in this signal region with a combination of Monte Carlo simulations and extrapolation from outside this region (data-driven backgrounds). This obviously has both statistical and systematic uncertainties, and a lot of the hard work is in getting this right. For instance, you might use independent Monte Carlo generators, or define control regions to check your estimation in.
- "Open the box" and look at the signal region data. If you see an excess over the estimated background, calculate the significance. Even here, if you were looking for a new particle, for instance, and the excess isn't a localized bump, this would be an indication that the background estimation may be flawed.
- If the significance is over 5 sigma, you have a discovery.
As you can see, you aren't using a null hypothesis that you already know is flawed. Your null is specifically "no signal exists". A positive deviation that isn't a statistical fluctuation is by definition a discovery of a signal.
Things like malfunctioning equipment go into the uncertainties if they occur in a way that the researchers have considered (which often means uncertainties are set conservatively if the equipment behaviour is poorly understood). If they occur in a way that no one considered, there's no statistical trickery that's every going to compensate for that. If we just went straight to setting limits, we would still see a deviation from what's expected there if equipment malfunctioned in a manner that faked a signal.
> In summary: If there is anything wrong with the null/background model at all, the p-value will converge on zero, it is just a matter of collecting enough data.
In summary: Assuming the background estimation is correct, the p-value will only converge to zero if a true signal exists. If the background estimation is wrong, then the background estimation is wrong and this is a problem with the background estimation, not hypothesis testing.
I'm by no means claiming the system is perfect, but it isn't systematically flawed the way you seem to be claiming.