First off, it's important to realize that the above approach is technically correct. The odds of getting 7 zeros in the hash of your research paper by chance are indeed less than 1%.
However, this is obviously a useless test, its probability of rejecting the null hypothesis given that it is false is likely 1/128... We don't really expect the phenomenom that we're measuring to produce good hash preimages. In technical terms, the test has no power.
However, to even talk about the "power" of a test, one needs a model of what the effect looks like. For instance, we expect the medicine will increase the rate of cures, we expect education will increase job success, etc.
And yet, the entire freaking point of p-value testing is to maintain a pretense that Popperianism makes sense. Validating a hypothesis? While my simple minded fellow, that would be anathema! Don't you know that Popper explained that you can only refute hypotheses! So we'll refute the null hypothesis instead.
Fine, but if you do that, the hash-function method works. The reason it sounds stupid is that Popperianism is wrong.
The focus is on rejecting false hypotheses, not on proving correct ones.
Translated into statistics, this gives you p-values. To pretend they're not relying on induction, people test "the null hypothesis", hopping to statistically disprove it.
My point is that, in the absence of another hypothesis, there are an infinite number of ways you can disprove the null hypothesis. In real science, the actual hypothesis is always lurking in the closet, so why not actually model it and compute likelihood ratio?
There is actually a much more rigorous framework for epistemology, it's algorithmic complexity. While it's not practical to use (the most common version is uncomputable), it's an ideal that helps thinking about the scientific method.
Popper himself doesn't say that disproving the null hypothesis leaves you with the correct one. The problem lies with people who use disproving the null hypothesis as a way of validating the alternative hypothesis.
The fact that they feel compelled to seek indirect validation of the model via disproving one other hypothesis is closer to logical positivism, isn't it, where people feel compelled to have a validation for every statement?
So isn't p-value testing just logical positivism in the garb of Popperianism?
Popper himself in this case would have tried to disprove the actual hypothesis and acknowledged it as a probably valid one in case if didn't find any mistakes, wouldn't he?
Anyway, Popper believed that you could never really give affirmative support to a hypothesis, only fail to falsify it. A good example of why he believed this is that, after all, Newtonian gravitational mechanics passed hundreds of experimental tests for centuries at a time, but was still not actually as true as Einstein's later Theory of General Relativity. Popper also held that the actual process of coming up with theories required a Real Human Thinker somewhere in the mix, and that the experimental procedure of forming hypotheses, testing them, and converging to truth required the magic secret sauce of Real Human Thought.
This combined very well with Fisher and Neyman-Pearson statistical testing, under which one infers, "the conditional probability (likelihood) of my evidence given absolute belief in my null hypothesis is very small", and thus treats the null hypothesis as falsified in Popperian fashion. Of course, in actual statistical testing, the critical values of the test statistic will take a particular alternate hypothesis (parameter value or distribution) into account, and thus a very low likelihood of the evidence, given the null hypothesis, is taken (note the italics: this isn't really a valid probabilistic inference) as support for the built-in alternative hypothesis.
This is, of course, complete bunk: we've done no quantitative reasoning whatsoever about the hypotheses themselves, null or alternative. Which leaves us with exactly the same problem in frequentist statistics as in Popperian philosophy of science: there's a big box marked "And then the magic of scientific thought happens, in the magical mind of an actual scientist!" Well, we might ask, if the philosophy of science is supposed to guide scientists as a genuine epistemological tool, what does Popper advise that the scientist should think and believe? And the answer is: conjecture things and test them, hoping that intuition and logic will guide the scientist to conjecture things which later turn out easy to test but difficult to falsify, this being taken as approximating eventual truth.
This gets rid of a problem that only exists in the academic study of epistemology: the Problem of Induction. But look what a price we've paid to sidestep it!
Scientific epistemology as Bayesian inference, on the other hand, has considerably fewer such problems. The two jobs remaining for the scientist are to invent coherent, testable hypotheses and to assign priors to them. Testability is given a rigorous meaning: a hypothesis is testable when it generates a non-uniform likelihood distribution over evidence. Priors are grounded in subjectivity, which sounds dirty but only really matters at small sample sizes (where frequentist inference would be weak anyway). Inductive inference is then given a rigorous meaning: model the set of hypotheses as a set of mutually exclusive propositions yielding nonuniform likelihoods over the evidence, collect data, and then use Bayes' Rule to move information from your (assumed) belief in the data to your (malleable) belief in the various hypotheses. This is rigorous probabilistic inference, since Bayes' Theorem is derivable directly from the definitions of conditional and joint probability.
(Philosophically speaking, this is also how you can escape the trap of trusting only in a priori reasoning: Bayes' Theorem and Cox's Theorem give a solid a priori argument that you will always "lose at life" if you don't reason using evidence and probability theory, compared to someone who does, therefore you really, actually should and we're not just making this up.)
And then we can do what /u/murbard2 is alluding to, and go for "hardcore mode" on Bayesianism: Objective Informative Bayesianism (better referred to as Algorithmic Bayesianism). This is the circumstance in which we explicitly treat statistical inference as a way of moving information from belief in data to belief in hypotheses (Bayesian) or from belief in hypotheses to belief in data (frequentist), treat models/hypotheses as computational objects, and treat probabilities as measuring belief in terms of information (usually measured in "decibels" if you're being casual or "bits" if you're a Real Information Theorist). Or, in fact, this approach lets us dissolve the very notion of belief: information/evidence becomes something that weighs upon a hypothesis space to locate the truth, like how mass weighs upon space-time to create gravity. You can thus view the real physical system under examination as emitting information when an experiment takes place, with some anti-information (randomness) mucking it up slightly. Your inductive process starts with a hypothesis space "weighed up" or "evened out" with anti-information (ignorance) which then comes to reflect the real world as more and more space is "weighed down" by evidence-information the world emitted. The primary problem, then, is how to assign priors: how to decide how much information must "weigh down" a particular hypothesis before it becomes "heavy" (believable with confidence).
Notably, when we phrase it this way, probabilistic phrasings of Occam's Razor become intuitive to the point of obviousness: a simpler theory will have a higher prior, meaning we need less information to make us confident in it. Thus, given a set of theories which all explain exactly the same data exactly as well (generate the same likelihoods for that data), and an assignment of priors such that simpler theories have greater priors, the theory about which we are most informed by the evidence, and in which we can be most confident, must therefore be the simplest.
(Of course, you could assign perverse priors that order your theories from the most complex to the simplest (in descending order of belief), but this just means you will require more evidence to arrive to the same answers.)
Algorithmic information theory then gives ways to formalize Occam's Razor by assigning prior probabilities to all possible Turing Machines based on their algorithmic complexity. This isn't very useful in real life, but does actually solve the Problem of Induction and give a provably optimal way to generate predictive probabilities for anything computable at all (ie: just about anything).
This sounds very interesting but there are no hits on these terms other than your post. Any references?
But that doesn't answer my original question. The problem here being referred to was people disproving the null hypothesis to indirectly validate the alternative hypothesis, right?
Is that still Popperianism? Because as far as I understand it, Popper only tells to keep trying to disprove the actual hypothesis/hypotheses.
Take the case of string theory, it may not ever be possible to have direct validation of string theory, and it still turns out to be mathematically useful. As long as there is no evidence that string theory is wrong, Popper tells that people can still use it where it is mathematically useful. According to you, I am guessing mathematically useful theories will have a higher prior. Both these are in opposition to logical positivism, where people hold that string theory isn't meaningful because it isn't validated.
Of course, I accept your statement that Popperianism isn't the correct solution to the philosophical problem of how to guide scientists towards the truth. As far as I understand it, its just a way of saying that things can still be meaningful if it hasn't been validated yet.
But people misusing p-value testing as a way of confirming the actual hypothesis without testing the actual hypothesis is just misusing Popper's idea as an indirect way to have some validation for their actual hypothesis, and that seemed(seems?) to be closer to logical positivism to me because of how they feel compelled to have some validation to say that their statement is meaningful.
I actually did answer, but it was kind of buried. Yes, using p-values to disprove the null hypothesis is Popperianism, at least insofar as it holds that you can't ever really provide quantitative validation to your alternative hypothesis and have to sort of wave your hands at its non-falsification as if that meant something. It seems almost but not quite like positivism because Popperians desperately want to validate things but are committed to an ideology telling them it's categorically impossible to do so. So they conjecture really fervently instead.
If this all sounds kinda stupid, well, I did give a whole rant on Bayesianism, which does let us quantify validation of hypotheses directly.
Thanks for continuing to answer. I am learning new stuff.
That's a great way to put it, I like it a lot. How do you want to validate your hypothesis then? I admit that it's quite a thorny issue for me. I like Box's quote "All models are wrong, but some are useful" --- I assume that if I poke at my model long enough, I'll always be able to falsify it eventually (eg. show that the error term isn't Gaussian). Nevertheless, a model can still embody enough truth to be useful in a given context.