Psychology study that induced "reproducibility crisis" was wrong: researchers
slate.com
slate.com
http://neuroneurotic.net/2016/03/03/the-non-replicability-of...
From where I’m standing, social and other forms of traditional psychology can’t say the same. Small contextual or methodological differences can quite likely skew the results because the mind is a damn complex thing. For that reason alone, we should expect psychology to have low replicability and the effect sizes should be pretty small (i.e. smaller than what is common in the literature) because they will always be diluted by a multitude of independent factors. Perhaps more than any other field, psychology can benefit from preregistering experimental protocols to delineate the exploratory garden-path from hypothesis-driven confirmatory results.
I agree that a direct replication of a contextually dependent effect in a different country and at a different time makes little sense but that is no excuse. If you just say that the effects are so context-specific it is difficult to replicate them, you are bound to end up chasing lots of phantoms. And that isn’t science – not even a “soft” one.
https://web.archive.org/web/*/http://wjh.harvard.edu/~jmitch...
(I agree much more with your link).
The article then goes on to add the the meta-study was awful in it's study selection, giving the example: OSC researchers tried to reproduce an American study that dealt with Stanford University students’ attitudes toward affirmative action policies by using Dutch students at the University of Amsterdam.
The article then goes on to say that reproducibility is still a problem in science, but that the meta-study was simply terrible in it's own methodology. Low reproducability is less of an issue given the context of the problem space, the article says; the meta-study was wrong, both methodologically and in conclusion ('low reproducability' != 'unethical' or 'crisis').
I recommend reading the fine article.
If they do conclude things on "students" based on this study, then reproducing it on Dutch students is perfectly fine. If they did conclude something on "Stanford University students", then it is a bad reproduction. You see, there's a very good justification there, that you can interpret as subtyping, or the liskov substitution principle.
>'low reproducability' != 'unethical'
This is akin to saying "ok there's this theorem, but I won't give proof, just a bogus heuristic argument", or even worse, saying "that person is guilty but I won't justify why". Both these behaviors are unacceptable. Why would "social sciences" be above that?
> or 'crisis'
A field is filled of politically oriented baseless claims, but this is not a crisis, just another day in Oceania.
That said, if the meta-study was sloppy, then it still justifies that research standards should have been upped a while ago; the "terrible"" meta-study was accepted precisely because terrible papers are accepted in the first place.
No. Just fucking 'no'. It was a study on affirmative action policies, repeated in a distant country with different public opinions on welfare and social fairness. If you think that that's fine, then you really have no claim to be declaring what psychology should and shouldn't be.
> filled of politically oriented baseless claims
Fuck me, but it's funny when detractors claim that psychology is full of unscientific behaviour, while they're way overgeneralising the entire field on a subsection of it.
You could also try to answer to the arguments of harry8, who gave you polite and well-written points, that you consciously ignored and resorted to name-calling ("you're heavily clueless about what psychology is") when you have absolutely no element to back your claim.
>Fuck me, but it's funny when detractors claim that psychology is full of unscientific behaviour, while they're way overgeneralising the entire field on a subsection of it.
What I mean by filled is that this subsection (which exists in any field, even mathematics) is an important fraction of this particular field. And before you bring this up: I read a lot of important papers in the field, and they were all below my expectations, which were admittedly not very high. I came to think that the "journal of irreproducible results" is one the most serious publications in soft sciences.
>Apparently arguments become stronger when you add "fuck" qualifiers
Fuck me!
https://hardsci.wordpress.com/2016/03/03/evaluating-a-new-cr...
Most of the popular books on learning now suffer from this, emphasizing strategies that work fine for short-term rote memorization tasks (like memorizing random word lists), but have little impact on or can even hurt higher-order learning and understanding.
The 'researchers' chose to do a exclusive pre-release with a journalist/media before you could check the data says it all, they are not scientists.
The more I read about Psychology the more I think it really really isn't currently a real science.
https://xkcd.com/435/ - Currently I think it starts to cut off around chemistry to be honest.
They are saying there is no reproducibility problem, because it's perfectly normal to be able reproduce just 39% of studies?
Whatever the actual error rates turn out to be, something like that is going to cut pretty deeply into measured reproducibility rates, isn't it?
On the other hand, if probabilities of methodological flaws are low for reproductions, then the multiplicative effects of the original method and the reproduction method are dominated by the original method, in such a way that more replications means multiplying probabilities together on the replication side, which (if flaw rates are less than 50%) goes to zero pretty quickly.
Meaning that consistency among reproductions is strong evidence regarding the validity of the original, and so the replications are worth the cost.
There's a significant qualitative difference between cases where replications are likely to be flawed vs. cases where they aren't. Cases where different replications are likely to have correlated flaws (even when flaws overall are unlikely) would also be problematic, because different replications might not give you additional information.
I think it matters to be very clear about all this, and not give any impressions that could be easily misconstrued to cast doubt on the utility of replications. Making some back of the envelope estimates using a 50% flaw rate strikes me as a dangerous way to talk about it (unless we have evidence that the flaw rate really is that high).
I have exactly zero idea what the actual ratios are. But the actual ratios absolutely do not matter for the purposes of illustrating how the methodological errors compound.
If the flaw rate is >= 50% and the number of replications is low, it presents a case where it's of little social value to do replications, depending on their cost (which is usually high).
From your original comment, the way a naive reader will view it is that replications often don't work, because you primed them with a number you just pulled out of nowhere (50%). That will matter much more than the other detail about it being multiplicative, which doesn't really matter much at all no matter what the flaw rates are, because it all hinges on the flaw rates of the replications and how many replications you do anyway.
I know that you did not mean to misrepresent it with the 50% number. I'm just harping on this because by happening to use a number like 50% it has the potential to distract and do more harm than good, regardless of anything else in your comment or even whether or not the 50% had anything to do with your main point.
Sounds like a situation that could lead to a millennium of stagnation, similar to geocentrism.
Try building a bridge with that.
Social Science methodology is piss poor and needs to improve. People don't take it seriously for a reason.
Strictly on the basis of significance — a statistical measure of how likely it is that a result did not occur by chance — 35 of the studies held up, and 62 did not. (Three were excluded because their significance was not clear.) The overall “effect size,” a measure of the strength of a finding, dropped by about half across all of the studies. Yet very few of the redone studies contradicted the original ones; their results were simply weaker.
Also: The only factor that did [affect the likelihood of successful reproduction] was the strength of the original effect — that is, the most robust findings tended to remain easily detectable, if not necessarily as strong.
Here's the HN discussion of the study:
Nope!
"If all 100 of the original studies examined by OSC had reported true effects, then sampling error alone should cause 5% of the replication studies to “fail” by producing results that fall outside the 95% confidence interval of the original study"
It'd have to be a pretty poor(read corrupt) field of study of all the random-ish studies chosen at close to 95%?
If the original studies are legit, than many/most of them would be at much higher confidence intervals?
Only scam fields would they all be around 95%?
https://xkcd.com/882/ - this is more about how bad science creeps in legitimate fields, the authors would be implying all the studies in this fields are this bad.
This would be my reading, not sure....
95% confidence is just about the lowest that a scientist in any field would think about publishing with, but in many fields- especially those whose subjects are the delicate and high-variance animals known as humans- statistical confidence is expensive. It wouldn't surprise me that lots of papers are put out that just barely reach the lowest standards of evidence, because grant-funded researchers couldn't afford the 1000 more test subjects it would take to get another sigma.
Which isn't great, but strictly speaking its not doing society a disservice, either- even uncertain knowledge decreases the entropy of our vast ignorance. It does, however, add an important dimension to how these studies should be interpreted- that small chance that any given study is incorrect matters, particularly when there's such a mind-bogglingly huge amount of research being done in the modern day. Its entirely unsurprising to read about a study claiming with p < 0.01 that a glass of wine a day will make your hair fall out, when you consider that there's been many thousands of papers published about things like wine. Confidence- and p-values alone aren't going to save us there.
Most studies would chose a 95% confidence interval so their statement stands.
Not being able to reproduce a political questionnaire done of Harvard students on U of A students should surprise exactly zero people, and should never have been used for verifying reproducibility.
I'll bet 100% the paper was generalising from Harvard students to some wider population, or why would you publish it. If the paper says "this thing we're publishing is only applies to Harvard students and has zero wider significance." I'll eat humble pie and apologise.
Psychology has a big issue of claiming things as fact that are no such thing. It has had this issue since forever. The history of psychology research is truly ethically hideous both in terms of what it did to people in experiments and what it claimed to have found that was then used to maim people in the general population. Research lobotomy if you have any doubts at all about that statement. Some fraud got a Nobel prize for research that he used to advocate "safe lobotomy with no adverse affects." (But maybe he fooled himself and it wasn't blatant fraud. Bah.)
These are the screams from the psychology establishment as they get called out for claiming so much that is just totally false. It's really hard to know something. There is an immense amount of work you have to do and you all you might find after decades is "dead end". Not in Psych research papers and they absolutely deserve the contempt they get for it.
So, you've pretty much publicly announced that you're heavily clueless about what psychology is, and are buying into that armchair critic opinion that it's a discredited arm of psychiatry.
Discussions on HN constantly revolve around psychology. How to negotiate for a better salary? Psychology. A/B testing? Psychology. How to give or take an interview effectively? Psychology. UX design and testing? Psychology. How to motivate yourself or others, or avoid procrastination? Psychology. Game theory? Psychology. And this is just the common discussions on the rather niche forum of HN. As I write this comment, the #1 article is "Dsxyliea", which is... psychology.
When you abuse psychology because of a brief, discredited and niche episode in its history, you're saying the equivalent of "web browsers are shit software, because IE on XP can't do SNI".
[0] http://alexanderetz.com/2016/02/26/the-bayesian-rpp-take-2/ and http://journals.plos.org/plosone/article?id=10.1371/journal....
Published, peer reviewed, cited work that can't be replicated is a problem. If it were a random sample of papers it is also a systemic problem.
Take your pick, it's a systemic problem that needs to be addressed either way.
"We constructed a sampling frame and selection process to minimize selection biases and maximize generalizability of the accumulated evidence. Simultaneously, to maintain high quality, within this sampling frame we matched individual replication projects with teams that had relevant interests and expertise. We pursued a quasi-random sample by defining the sampling frame as 2008 articles of three important psychology journals: Psychological Science (PSCI), Journal of Personality and Social Psychology (JPSP), and Journal of Experimental Psychology: Learning, Memory, and Cognition (JEP: LMC)...The first replication teams could select from a pool of the first 20 articles from each journal, starting with the first article published in the first 2008 issue. Project coordinators facilitated matching articles with replication teams by interests and expertise until the remaining articles were difficult to match. If there were still interested teams, then another 10 articles from one or more of the three journals were made available from the sampling frame. Further, project coordinators actively recruited teams from the community with relevant experience for particular articles."
Typical example would be premature generalisation of results.