HNHacker News
TopNewBestAskShowJobs

yurolis

12 karma · joined December 1, 2018

submissionscomments
yurolis··on Psychology’s Replication Crisis Is Real
I'll admit that might have something to do with it, although psychology GREs are misleading because a lot, probably the vast most majority, of those test-takers are hoping to go into clinical psych, especially to practice therapy. This leads to issues in a number of ways: (1) many of those applicants really don't understand what psychology is about, (2) most (90% in research-oriented programs) are rejected, and (3) among what's left, it has to be said that they probably don't need to be the most quantitative to be a good therapist, if they're not doing research.

Psychology is a strange science in that it's a mixture of people with very unquantitative backgrounds, and those who deal with very complex math and statistics. What many don't realize is that meta-analysis itself really was developed as a method in psychology (even if it technically has its origins earlier with Pearson). This registered replication work is an extension of that, again being done by psychologists. It's probably safe to say that more empirical and statistical research on the scientific process itself has been done by psychologists (along with statisticians and many public health researchers) than any other discipline.

In any event, replication problems happen in other domains as well. This has been documented empirically. It might be worse in the biomedical domain than, say, physics or chemistry, but it's not limited to psychologists. What I see in the neurosciences per se is just as bad, if not worse (because it's ignored more).

I think ignorance of issues pertaining to overfitting, etc. definitely contributes, but I also think that ignorance is pretty widespread, and the problems can be sort of pernicious in that they don't always operate intuitively.

yurolis··on Psychology’s Replication Crisis Is Real
I think the fact we're even discussing a "replication" crisis points to some small movement in academic research to address this (some journals are changing / have changed policies as a result).

There's probably different causes at different times. Some of these effects that are the target of replication tests are relatively older, when people were less aware of some of this phenomena. To be fair, some of this stuff is unintuitive: for example, some of the major journals would require internal replications, over several samples, the rationale being that if someone shows an effect in 5 samples with slightly different designs, it's probably "real." It's not like some of these things were just based on single samples (although some of them certainly were). Of course, now people are aware that 5 small samples does not large-sample replication make.

My intuition is that another cause is that academics is flooded with researchers collecting a lot of data, under a lot of pressure to produce positive findings to attract money. Hype, TEDtalks, grants, hypercompetition, and so forth. Academics now is horribly incentivized to be popular and bring in money (from peers, it's important to note), rather than to be correct. Add in a complex subject, like human behavior, and it's a recipe for disaster. Academics is also full of conventions, that attain the strength of power structures; when these are baked in it's hard to change them because you have people who attained power under those conventions in charge (nothing conspiratorial really, just people have their biases and blindspots).

yurolis··on Psychology’s Replication Crisis Is Real
I do research in psychology and also a little in machine learning (classification using text samples). I've actually been working with the Many Labs data (the project that's the focus of the article).

My impressions of machine learning match up with yours. There's just so many parameters that there's so many opportunities to capitalize on chance, combined with [necessarily] huge datasets that preclude replication, that it's hard to avoid. I have also been surprised at how much tweaking there is of parameters; I'm used to work in traditional statistics where more is derived from theoretical principles, as opposed to trying a bunch of values to see what "works." This all lends itself to overfitting. I strongly believe that a lot of the adversarial input work is basically capitalizing on this overfitting.

To be fair, I don't think people are necessarily being nefarious, I think people across all fields just don't appreciate the dangers of overfitting.

The one upside of this Many Labs work, mentioned in the paper, is that it tends to show that the common criticism of "how can you generalize from this sample of undergrads" is not so much of a problem. Not that it's not an issue at all, but if you're studying some basic cognitive process it's probably not going to matter that much if you use undergrads versus some perfectly representative sample of the population. People have shown this in different ways before but it's useful to know more about. Obviously with some things sociogeographic variation will matter more though.

One thing that might not be totally obvious from this article and others is that although some effects clearly replicate, and others do not, there are some effects that seem to be in a grey area, where the effects are probably real but much smaller than originally reported. The distribution of effect size estimate distributions is more continuous, with shades of grey, than these news reports would have you think. Whether or not it matters that an effect is tiny versus zero might not practically matter, but at some level it is important to be mindful of.