How to Tell Good Studies from Bad? Bet on Them
fivethirtyeight.com
fivethirtyeight.com
This article says that the prediction market correctly predicted 71% of the replication results of 44 studies, or 31 correct.
Assume the studies have a 50% chance of being replicable. Then a random coin would predict a mean of 22 correct with a std dev of sqrt(0.5 * 0.5 * 44) = 3.3. This sample has a z score of 2.72, which means there's a probability of 0.003264 (0.3%) of the random chance approach being correct 71% or better. So the result seems pretty significant. (Changing the assumed 50% to other values makes the probability even more extreme.)
Is their study of "using betting to accurately predict whether studies are reproducible" also reproducible?
A very interesting question about stats and studies.
Even the p-value of 0.01 seems SUPER strong, but in reality, falls apart in many cases.
Let's hypothetically suppose that we knew that 75% of the studies were replicable. Can we make a better coin flip prediction? If you had a coin flip that says yes 75% of the time, it isn't necessarily correct at a rate of 75% -- instead it'd be right .75^2+.25^2 = 62.5% of the time. In fact a coin that just predicted "replicable" every time in this scenario would be right 75% of the time. So I think maybe my null hypothesis should've been based on "did they do better than a parrot that always says yes", not a coin flip.
I think the math in the original problem stays the same, it's just you change it to "a coin that always predicts it will be replicable". And in that case, if the underlying rate of replicable was 71%, then their prediction market only does as well as the always-yes coin and is in fact not very useful.
The second is, how do scientists know which studies are good? This is a harder job than for you¸ because you're only ever made aware of papers that made it through the publication gauntlet. (No, they won't publish "anything" these days, not in a journal that gets any coverage.) For scientists, the task is harder, so prediction markets might be the tool they need.
It certainly isn't a fool-proof method of increasing accuracy, and it does favor popularity of a theory over other factors, but overall it's probably a nice layer of data to consider adding to the mix.
Here's a book all about it: http://www.amazon.com/The-Wisdom-Crowds-James-Surowiecki/dp/...
Yes, this is the critical piece. The results of the Reproducibility Project were not remotely a surprise to Bayesian observers. People like Gelman have been pointing out for ages (and I mean back to the 1960s) that the prior probabilities in these fields is low and necessarily a lot of the results were false positives. With the rise of meta-analyses, it is possible to have informative priors for particular fields of psychology or for psychology as a whole, which would let you make much better predictions about whether a result was real. But you can't use these in papers - authors are heavily biased towards using procedures or flat priors which are uninterpretable or grossly overestimate the evidence, and if you try to use any of the informative priors or more advanced models, they'll nag you to death with a thousand objections and complain about double standards and subjectivity and how this time is different and (ironically) bias. So for the most part, there's not much to gain in academic research.
But in a prediction market, you don't have to listen to the self-serving excuses or explain your reasoning, and there's something to make it worth your while.
[1]and of course it hasn't been reproduced yet ;-)
[Edited as enlightenment dawns ...]
Science advances one funeral at a time. The stock market accelerates the process by separating fools from their money.
This betting market only covers reproducibility, not truth. But since most reproducible findings are still false, betting contrary to bias is unlikely to work.
Toy example to illustrate my claim: Suppose a given finding has a probability p of replicating, but the biased market estimates q < p. This means that you must spend $q and if the experiment replicates you'll earn $1.00.
On average you'll win $p from these bets - $1.00 exactly p of the time. Your net winnings will be $p-$q. As long as your theory more accurately predicts the replication probability, you make money.
At least in the context of the phrase 'science advances on funeral at a time', usually folks are talking about studies that are wrong for reasons other than replicability. (Because in general it's pretty easy to convince folks that something is false if it can't be replicated, but much more difficult to convince folks that there is some deeper methodological or epistemological issue.)
I.e., observe a bernoulli trial and see 60 successes in 100 trials. Plan a followup study which will repeat the bernoulli trial 100x - predictions would be on the # of successes in the followup, not some statistical analysis.
The phrase "science advances one funeral at a time" is somewhat orthogonal to this idea - it tends to be about older scientists being unwilling to accept new theories in spite of repeated successes. Such scientists would lose their money if they bet or be ignored if they didn't.
Attempting to replace the role of independent replication with opinion is an awful idea. And that is the goal here, to replace, not to supplement. Of course, it doesn't seem replication attempts are very common in this area to begin with. So this is actually going to be a justification for continuing that pseudoscientific practice and to avoid checking all the previous results.
Suppose there is a parameter - say the probability q of a Bernoulli random variable being true. Based on your past experiments you have a posterior p(q). Then given p(q), you can easily compute a probability distribution on S(N), where S(N) is the number of successes you'd get from a bernoulli random variable with N attempts.
In code terms, you are just computing posterior.flatMap(q => Bernoulli(N,q)). (Using the inherent monadic structure of probability.)
This actually works in general. If you want to predict the outcome of a later experiment (the replication) you just compute posterior.flatMap(parameter => generateResult(parameter)).
Then you can simulate the results of future experiments with _n_ datapoints: sample 1 effect size from that distribution, generate _n_ datapoints assuming that effect size, run the hypothesis-testing, and return the _p_-value.
The fraction of _p_<=0.05 is your best forecast of whether the future experiment will succeed in reproducing it or not.
If you successfully replicate, then it has survived the challenge. The more challenges it survives, the more confidence we can have that the result is valid.
That said, we do not try to make replication challenging. We try to present experiments in a way that makes replication as easy as possible to perform. Exactly so that people don't have to take what we said on faith.
The point of the prediction market here is not to replace independent replication with opinion. It is to ensure that the energy that gets spent on replication is more likely to be spent effectively.
Ideally, of course, targeting replication efforts more effectively will increase the value of time spent in attempting replication. This should therefore increase how much effort is spent on replication. Which is exactly the opposite of replacing independent replication with opinion!
A lone report of some observation should never be believed. It should be verified by others who retrace the steps.
>"If you successfully replicate, then it has survived the challenge. The more challenges it survives, the more confidence we can have that the result is valid."
If the observation can be independently replicated, it shows that the methods required are understood well enough to communicate and that it is stable in the face of unknown influences. This increases our confidence that we understand what is going on and that the phenomenon is worth theorizing about.
It has nothing to do with an observation being true or valid. The observation was made, it happened. It is true. It is valid. (Excepting outright fraud, which is treated equivalently to some severely unstable phenomenon)
>"The point of the prediction market here is not to replace independent replication with opinion."
They explicitly say that is the goal in the paper. I like this method of eliciting priors, not that the goal is to substitute it for actual replications.
>"It is to ensure that the energy that gets spent on replication is more likely to be spent effectively."
If a study isn't worth replicating, then it wasn't worth doing and reporting in the first place.
See http://www.the-scientist.com/?articles.view/articleNo/43875/... for horrible replication rates in psychology. See http://journals.plos.org/plosmedicine/article?id=10.1371/jou... for evidence that medicine is no better. See http://www.nature.com/nature/journal/v483/n7391/full/483531a... for evidence that the research on which we base clinical drug studies is also in bad shape.
At this point anything that focuses attention on, "We should replicate SOMETHING" is a huge improvement. Eventually we should get to, "You should expect to be replicated." But we are a long, loong, loooong ways away from that now.
I am well aware of this problem. I expect you will have no more success convincing these people that replication is necessary than you would convincing a fervent religious believer their deity does not exist.
These juvenile research practices really need to stop though. It is driving the most intelligent and competent potential contributors away from careers in these areas.
This. Null hypothesis testing: Just stop it. We've known this is a bad idea for decades
This is more good data that gives me more faith in the www.Augur.net decentralized prediction market premise.
[1] https://www.cia.gov/library/center-for-the-study-of-intellig...
Minimally though we're looking to establish an ongoing dialogue between the different levels of an organization. Our belief is that people on the ground building product, interfacing with clients, etc. aren't consulted nearly enough about predictions that inform big strategic decisions. Instead, leaders are making decisions based on input from a limited number of SME's, data analytics, and their own beliefs. None of these are bad per se, but not leveraging your own people we believe is a huge opportunity lost.
Happy to follow up live/over email if you'd like. adam at cultivatelabs
So what do they mean when they say it correctly predicted the outcome? Are they just saying the odds fell on the same side as the reproduction indicated?
If so, that seems arbitrary. If the cutoff for a p-value is 0.05, then shouldn't we say that any contract selling for less than $95 predicts a reproduction failure?
Is 0.01 that low for such a crazy finding? Let's say you believe that it has a probability of 1 in 10,000. And that result really seemed really really unlikely. 1 in 10,000 might be generous. Then, after this study, the probability that it's true is 1%.