Or are they trying to say that cherry-picking is still a risk, but is less catastrophic than it is with a frequentist approach?
Or are they trying to say that cherry-picking is still a risk, but is less catastrophic than it is with a frequentist approach?
> Bayesian: "Shrug," I say. You can't mislead me by telling me what a real coin actually did.
> Scientist: I'm asking you what happens if I keep flipping the coin, checking the likelihood each time, until I see that the current statistics favor my pet theory, and then I stop.
> Bayesian: As a pure idealist seduced by the seductively pure idealism of probability theory, I say that so long as you present me with the true data, all I can and should do is update in the way Bayes' theorem says I should.
> Scientist: Seriously.
> Bayesian: I am serious.
> Scientist: So it doesn't bother you if I keep checking the likelihood ratio and continuing to flip the coin until I can convince you of anything I want.
> Bayesian: Go ahead and try it.
Maybe the part I’m struggling with is “You can't mislead me by telling me what a real coin actually did”. Is the author overstating the benefits, or are they trying to communicate something that I’m just not picking up?
I’d expect that frequentists are exactly as immune to this problem on a per-study basis, the only difference being that Bayesian statistics provide the ability to aggregate likelihoods from multiple studies in the hopes of counteracting bias.
This is the key line. This is the only type of cherry picking they are referring to: you run the experiment until the results look favorable and then you stop.
With a frequentist analysis the p-value depends on why you stopped, because that changes the reference set that the outcome is calibrated against (the set of outcomes that are considered as-or-more-extreme). This is illustrated by the example of flipping a coin and getting HHHHHT at the beginning. If the experiment was "flip a coin 6 times and then stop", then HTHHHH is as extreme as HHHHHT. But if the experiment was "flip a coin until you get tails" then HTHHHH isn't a possible outcome (because you would have stopped at HT), so it's not in the as-or-more-extreme set.
This is often used by Bayesians as an argument against frequentist statistics. Personally I've never found it very compelling, since (1) the effect tends to disappear as sample sizes grow (due to likelihoods approaching normal distributions in the limit under weak conditions) and (2) experiments generally have a predetermined stopping condition and I don't particularly care what the analysis would have been if the stopping condition had been defined differently. It is however a philosophical point in favor of Bayesian analysis.
Yudkowsky then presents this as a counterargument from the frequentist character:
> Scientist: Okay, that last claim in particular strikes me as very suspicious. What happens if I want to persuade you that a coin is biased towards heads, so I keep flipping it until I randomly get to a point where there's a predominance of heads, and then choose to stop?
Basically, "OK, if the stopping rule doesn't matter for Bayesians, what prevents me from using the stopping rule to manipulate the experiment by choosing to stop when the evidence looks favorable?"
I consider this a bit of a straw man because I've never heard a frequentist actually use this argument. I have heard Bayesians present it as a frequentist argument many times. But probably at some point in the past one or more frequentists did use this as an argument.
The Bayesian response is "You can try, but you probably won't be able to get very strong evidence in favor of a false hypothesis." Which is true and fine, as long as all the data are fully and fairly presented. But it doesn't prevent me from accumulating data over multiple attempts (by selectively failing to report attempts that don't turn out well): it's unlikely that I'll be able to get a 20:1 likelihood ratio against a true hypothesis in a single run (this is what the discussion about python programs is about), but it's not hard to get say 2:1 likelyhood ratio against a true hypothesis in a single run, and then do that several times, discarding runs that don't work out for me.
Yudkowsky's analysis doesn't address this. He assumes throughout that the data are all fully and fairly presented; I'm not throwing out any data from days when it didn't work out in my favor or anything like that. He also (as I pointed out in another comment, and as levocardia illustrated wonderfully in yet another top level comment) ignores the many other ways of manipulating outcomes other than just deciding when to stop collecting data: choosing which variables to include/exclude, choosing to analyze only a subpopulation, choosing transformations to apply to variables, and many more choices that a data analyst can make to obtain more favorable results. Bayesian analysis is not immune to any of these other manipulations.
In any case he does mention it won't fix all the problems but that it can clean up some of them and make addressing other problems a bit more straightforward. With p-values, sometimes even honest experiments end up with misleading conclusions. It takes a decent amount of complex understanding to perform properly, while something like gathering data and doing a conditional probability calculation has presumably less risk of honest mistakes
His claim about likelihoods being immune to p-hacking is far too strong.
> With p-values, sometimes even honest experiments end up with misleading conclusions. It takes a decent amount of complex understanding to perform properly, while something like gathering data and doing a conditional probability calculation has presumably less risk of honest mistakes
Maybe. I've seen an experienced Bayesian statistician who had published a paper about Lindley's paradox fall prey to Lindley's paradox and publish a misleading conclusion as a result. Bayesian analysis also has some pitfalls. Name withheld to protect the innocent.
that since bayesian methods account for both the likelihood it is real and the likelihood it is not real, the observations are gaining evidence AND counter-evidence together. So as long as you don't lie about what a real coin did, then you are describing the actions of a real coin and therefor cannot mislead