We found only one-third of published psychology research is reliable – now what?
theconversation.com
theconversation.com
A good chunk of our social and educational policy is informed by psychology and sociology.
Since we have all grown up within various systems and institutions that have been largely informed by bad research, it's unlikely much will happen, since we see the claims made by these charlatans to be normal or common sense.
Nonetheless, it will be amusing watching the charlatans try to defend themselves in the coming years.
No graduate student is incentivized to re perform existing experiments to verify them, no they find much more success in creating a unique experiment which tests a previously untested hypothesis.
This is modern academic science. If you aren't doing something newsworthy, you're not getting any funding.
If you look at how this thing with psychology started ("replication bullies") or at what Brockman said about the LaCour scandal earlier this year (he didn't pursue publication initially because he thought he'd risk his career), it's not that they aren't incentivized, it's that they are actively discouraged.
Really, those researchers trying to do something novel but being incompetent at it, should put their skills to use replicating something else instead. Sure this would lead to them being seen as 2nd class scientists. But we still need that service, just as we still need cleaners even though they're seen as 2nd class workers.
Universities are the ones providing the incentives. Their management should be stopping and thinking about if their policies are adding value to scientific knowledge or subtracting it.
I'll take it a step further. The reason we're hearing about psychology and not sociology is that few people take sociology seriously, as a science.
Sociology is a science like political science is a science in that any irrefutable claim constitutes a valid theory. Evidence is secondary and the notion of falsifiability doesn't even come up.
Psychology has a very hard-science core, whose reputation is so tarnished by its "social" counterpart that no self-respecting researcher wants to be called a psychologist anymore. Replace your searches for psychology with cognitive science and you'll find a lot of good, reproducible, research such as this: http://cavlab.net/?lang=en
Instead, I think there are three problems:
1. The inherent problem of inference. When you set p < .05, then 5% of studies will yield false positives. This can be mitigated by a variety of statistical approaches, but it's an irreducible problem on a fundamental level.
2. Certain (sub)disciplines of science aren't very scientific. In particular, they suffer from a high incidence of unfalsifiable claims and hand-wavy definitions. A good example, I think, is the shame vs guilt literature.
3. Certain (sub)disciplines of science are highly politicized. I point my finger at anything involving "identity" (in the sense of "identity politics"), "diversity" (sex, age, ethnicity, sexual orientation, etc) or "bias".
Regarding point 3, take a look at the faculty and students of any university you'd like and take stock of how many Black people are working on racism, women on sexism, LGBT on sexual orientation, etc. As a cognitive neuroscientist, I don't want to claim that people can't study themselves, but in light of the serious reproducibility issues that social psychology is facing, I think this correlation is anything but trivial.
I strongly suspect that the "hard psych vs social psych" dichotomy can be further broken down, and that the sub-disciplines listed in point 3 account for a disproportionate quantity of false-positive results, on accounts of flawed methodology. I suspect that much of this research is actually an attempt to validate political opinion.
I may be wrong about all of this, but at least the core claim testable with a handful of linear regressions.
I'm not a statistician, but as I understand it, a problem with the p value is that the standard calculation hinges on assumptions that are not necessarily met by the data, and if those assumptions were correctly accounted for, p would be higher.
If p < 0.05, and in practice 50% of published results are bunk, then regardless of the underlying math, a practical rule of thumb seems to be: Multiply p by 10.
If more than 5% of your studies are failing to replicate, then something else is involved, but it's not necessarily fraudulent. Post-hoc hypotheses are usually the biggest culprit, but it's insidious to the point that researchers often don't realize they're doing it.
The distinction is important. Reproducible research is a lot easier but there are still famous examples of attempts to reproduce published results that revealed serious flaws. Cases where the data are kept secret to thwart reproducibility are thoroughly suspect. Replication is a harder problem, both in being able to perform a potentially costly and difficult experiment, and in the statistical analysis to show equivalent results.
I did not find this clear. What is the distinction you are trying to make?
Replicate: perform a new experiment with new subjects, apply the same methods, see if you get statistically equivalent results.
In many cases, they say a finding was not reproduced simply because the new study didn't quite meet the cutoff for "significance." However, if you do a meta-analysis of both studies, you find that the likelihood of the effect being real is higher, rather than lower.
In other cases, the replications were done poorly. The first step of many studies is calibration. You do tests to find out what sort of things your subjects are familiar with, and calibrate questions accordingly. A common mistake in the replications was to skip the calibration step, and simply reuse the questions on subjects with different backgrounds.
This publish-and-forget behavior annoys me about researchers. Sometimes I find results that look like they'll useful for my (non-academic) work, but the paper containing them is the final say on the matter. There's no version 2, there's no other work improving on it. It's just a dead end.
Stop talking about "psychology" like it's a monolithic thing. Start by separating "hard" psychology (perception, attention, psychophysics, etc) from social psychology.
Papers about saccadic adaptation are, I suspect, much more reproducible than those about implicit association.
That puts them in the general level of quality for all bio-medical research, so, yeah, deeper issues indeed, but I find it fascinating that this sub-group is so good.
The problem with social psychology, on the whole, is that it's highly politicized. This is especially true with regards to any study involving race, gender, or "identity" (in the sense of identity politics).
The boundaries are fuzzy, and you'll find people with psych, medical, neuro, math, CS and philosophy degrees in experimental cogsci labs.
I think your hypothesis conflates a few notions, for instance hierarchical organization and study design. I think the relevant factor in your comparison is not the level of analysis, but the measure. "Early life event X" is likely to be measured by a survey, while studying neuronal populations is in the realm of electrophysiology.
>"The first is a p-value, which estimates the probability that the result was arrived at purely by chance and is a false positive. (Technically, the p-value is the chance that the result, or a stronger result, would have occurred even when there was no real effect.) Generally, if a statistical test shows that the p-value is lower than 5%, the study’s results are considered “significant” – most likely due to actual effects."
No, there is endless literature on this. Just to start:
http://andrewgelman.com/2015/07/21/a-bad-definition-of-stati...
Taking this opportunity to plug an open source application that I work on, for cataloging and verifying independent replications, called Curate Science: https://www.curatescience.org/
If you read the post where that quote comes from [1] they make a number of points about how a better and more rigorous definition of replicated is needed, because p-value alone doesn't tell the whole story.
[1] http://simplystatistics.org/2015/10/20/we-need-a-statistical...
No, the p value is the probability of observing a result at least as extreme as your own given the null hypothesis is true. There are two errors here
1) Transposing the conditional: P(A|B) != P(B|A) http://rationalwiki.org/wiki/Confusion_of_the_inverse
2) Deviations from the null hypothesis can occur even the absence of a treatment effect, ie one of your model assumptions is wrong, baseline differences, etc.
Anyway, that is why instead of statistical significance researchers need to estimate the size of the effect. If you estimate the effect is in the range -1 to +2 then you publish that result. Others also estimate the range and see if these are consistent with each other.
My wording was poor but that's what I meant.
>"So should intersecting confidence intervals be our definition of replication? This too has a flaw since it favors imprecise studies with very large confidence intervals. If effect size is ignored, we may waste our time trying to replicate studies reporting practically meaningless findings."
I don't see why the definition of a replication should have anything to do with practical use or precision of the estimate. These are other important, but different, issues.
Intersecting estimates of the plausible range is fine as a definition. The observations are consistent with each other. The best way to make these estimates is yet another tangential issue.
Continue avoiding psychologists whenever possible.
http://www.federalreserve.gov/econresdata/feds/2015/files/20...
Carl Jung claimed one of the chief factors responsible for mass brainwashing is scientific rationality.[2] Society worships the Goddess of Reason while frowning down on "irrational" and non-verifiable religious testimony.
Now that science is proven to be systemically corrupt, what will "rational" people base their understanding on?
Aside: I designed a personality test / psychoanalytical tool that was inspired by Carl Jung. It's called Critical Stimulus and it can be found at https://www.thegamecrafter.com/games/critical-stimulus. The printable version can be downloaded at gumroad: https://gumroad.com/l/criticalstimulus
.
[1] http://www.economist.com/news/briefing/21588057-scientists-t...
[2] https://en.wikipedia.org/wiki/Self_in_Jungian_psychology
This is not a new notion in the least. The "rational" people will continue doing what they've always done: revising their conclusions.
Science is a process, and while we can question the notion of "scientific progress" in philosophical terms, this is neither a new idea nor evidence that science doesn't work. It's certainly not evidence that science is no better than irrational thinking.
>Carl Jung claimed one of the chief factors responsible for mass brainwashing is scientific rationality.
Few people take psychoanalysts seriously these days, in large part because of their long record of absurd claims and shoddy clinical work. Jungian theory has it's place in a conversation about literary theory, but not in a conversation about science.
I'm sure there are scientific aspects that can be brought to bear on a situation, but for a lot of people the "literary theory" part is just as helpful. Especially when you stop to think how much neurosis is fueled by pop culture (i.e. status envy).
And empirically speaking, psychoanalytic therapy has a piss-poor record in dealing with mental illness.
If you're looking for spiritual guidance, then maybe a psychoanalyst can help. If you're looking for clinical efficacy, they demonstrably don't.
I'll leave it as an exercise to the reader to make a relevant Google Scholar query, but be careful not to confuse psychoanalysis with clinical psychology.
> for a lot of people the "literary theory" part is just as helpful.
Fine, but this is a conversation about psychology as a science.
This is a debate in semantics: research that is not successfully reproduced is not scientific.
It should be part of that which society frowns down on as non-verifiable testimony.
Natural sciences or engineering-related research that cannot be reproduced is nothing but scientific theater. There are various reasons why people do that (including financial ones in the drug industry), but that's a discussion for another time.
This is not proof that "scientific rationality" is brainwashing people. If anything, it's proof that non-reproducible research dressed as legitimate science is dangerous -- so dangerous, in fact, that it can put lives in danger (e.g. when it happens in drug research).
That's both untrue and fallacious (see: no true Scotsman).
If you set your p-value threshold at .05, then one in twenty experiments will produce a false positive. As such, plenty of research is conducted in a benevolent and meticulous fashion, only to yield a non-reproducible result. It's still scientific; it's just not true.
I don't agree with subliminalzen, but (with all due respect -- really!) your comment is hogwash.
No it's not. A sound scientific approach requires that a theory be based on reproducible results. If an experiment that verifies your theory confirms your result today and infirms it tomorrow, then the theory, the experimental approach, or both, are wrong.
Of course, experiments that can't be reproduced are part of the scientific endeavour. Every discovery comes at the end of a long sequence of experiments with results scattered all over the grah. But treating them as anything other than stumbling steps that help you refine your understanding of the problem or as dead ends is as unscientific as it gets.
Which isn't at all what I'm suggesting.
I'm arguing against the idea that applying the scientific method and getting a false positive makes the effort unscientific.
So yes, it is both untrue and fallacious.
The article you are mentioning in [1] refers to the fact that a quarter of the published drug research is not successfully reproduced.
That doesn't mean that people published papers saying "Hey, we did this experiment. Its results cannot be consistently reproduced, so we think it's unconclusive/because our theory is flawed with regards to this or that/because the experiment was flawed with regards to this or that and we think it can be refined by changing this approach or that apparatus".
It means that a quarter of the published papers say "Hey, we did this experiment which offers conclusive proof of X", but it turns out that their experiments cannot be consistently reproduced, so they're proof of exactly nothing.
That is unscientific.
This is patently false. Publication is never a claim of conclusive proof; it's a claim of evidence.
I'm sorry, but you are wrong about this. False-positives don't suddenly make the experiment un-scientific. You're very misinformed about how science works:
- False positives are part of the landscape
- Contradictory evidence is part of the landscape
- The above issues are resolved by tracking reproducibility of results
You can come to a wrong conclusion using valid scientific means. The scientific method hinges on the assumption that research will eventually converge on a correct result.
I said they published papers in which they claimed they conclusively proved something, and it turned out they didn't conclusively prove anything. Specifically, because their results couldn't be reproduced.
In case you're not familiar with how experiments are carried out in natural sciences, "results couldn't be reproduced" means that
1. They claimed they got <these results> with p < <this threshold>
2. Some other guys repeated the same experiment ("repeated" as in they administered the same substances, to a sample of equal size under similar conditions and measured the same parameters under similar conditions) and it turned out that on their results, p was through the roof.
In some cases, that was simply because the authors didn't publish enough information for their experiments to be repeated (I was close to making that mistake, too. Thank God for review committees). But in most cases, that simply happened because authors cherry-picked data or "optimistically" interpreted results.
(Edit: Responsible review committees can sometimes spot the latter, but it's very hard to deal with the former. The correct thing to do is to have all researchers publish all their experimental data, even the one which wasn't included in the papers. A lot of researchers agree, but you'll find that a lot of companies that employ researchers actively invent reasons why their researchers shouldn't do that.)
> If you set your p-value threshold at .05, then one in twenty experiments will produce a false positive.
> The reason is simple: given a p-threshold of .05, one in five experiments will yield a false positive.
Make up your mind already.
Exactly where you typed It means that a quarter of the published papers say "Hey, we did this experiment which offers conclusive proof of X"
Again, this is patently false because they did not publish papers claiming conclusive proof. They published papers claiming evidence in favor of a theory.
>Make up your mind already.
There's no reason to be disrespectful over a mistake. I meant 1 in 20 (5%, hence the mix-up).
Returning to the point, it takes incredible mental gymnastics to argue that a false positive automatically degrades the status of a study from "scientific" to "unscientific":
1. The adjective "scientific" describes a method, not a result. Those speaking of "scientific results" are either (a) referring to "results of a scientific study" or (b) confused about what science is (namely: a method, not a result).
2. A false positive degrades the status of a result (not a study) from "evidence in favor of X" to "not evidence in favor of X".
I must respectfully insist that you are wrong.
No one said anything about A false positive!
"Cannot be reproduced" means there were a lot of false positives. So many, in fact, that you can't really draw any conclusion from the experiment. (Edit:) Or more to the point, that the p value the original authors claimed was bullshit.
Reproducing an experiment means reproducing both the experimental technique and the sample.
Yup, and now that the science in question was attempted to be verified and shown to be poor, we can improve the process. Try that with religion.
> Now that science is proven to be systemically corrupt
Huh? Anyone who took psychological research as gospel does so at their own risk and ignorance. We've long known psychology is a soft science little better than voodoo. For example, most of the psychological establishment still default to 12 step programs for addictions. Guess what? It works about as well as any faith based healing program.
This is just another measure of validating the scientific method. Which is to say we don't know everything but we're learning.
Can we?
To the scientists devoted to furthering knowledge, they will improve. But there will be many who hold onto this. A subset of what we currently see as science is probably better described as religion already and we are soon to see this distinction clearer than in the past.
If we couldn't, this HN discussion wouldn't exist.
What I saw, read, and heard as I began my climb up the Ivory Tower made me turn back and go into industry instead. Once you are inside you begin to see the corruption. From the lack of pay to people who were allowing their views to corrupt their work (often seen by tweaking definitions of terms, especially in the realms of sociology and social psychology).
I still want to go back, but I would like to do so once I'm personally financially secure as that would remove one of the vectors for corruption.
P.S. To be clear, nothing was bad in the lab I worked in, perhaps because there was little to make political about creating VR landscapes and interactions.
Oh yes, and all the scientific advancements will stop working from now on since science is proven to be corrupt. I can hear the satellites and the MRI machines crashing because science doesn't work anymore...
What this means is that studies are less rigorous than promoted to be.
The guy who developed PCR (basically DNA testing) also has some interesting views: https://en.wikipedia.org/wiki/Kary_Mullis
From this I conclude that scientific advancement doesn't depend so much on commonly accepted scientific claims.
Nothing, I just gave those examples of people to show science can be systematically corrupt/wrong (since it is in their minds) and there will still be advancements.