What’s the difference between Derek Jeter and preregistration?
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
I agree that yes, preregistration is not a panacea that will fix the reproducibility crisis or whatever.
But the problem that it is solving is not to improve the quality of studies that get published. It only does that implicitly. What it does is allows us to factor in the unpublished studies. It gives us access to previously unpublishable null results.
Would this have debunked the ESP research? No, that study would have gotten published because its results supported the expectations of the experimenters. BUT all the attempts to reproduce that would also have been pre-registered, and their lack of publishing (or the publishing of null results) would have effectively debunked the original.
The "ESP guy" is J.B. Rhine. Not everyone knows/remembers that ESP was considered scientifically verified for a good couple 20th century decades, such that many people who considered themselves scientific thinkers felt they had to accept it's validity even though they had been suspicious. It was not just one study to be "debunked", it was literally a whole field (although granted one mostly controlled by Rhine, but that was not that unusual of scientific fields back then, not sure about now).
Makes you think about the whole endeavor of science. On the other hand, it was eventually abandoned (not with a big debunking bang so much as just kind of forgotten). One thing it does is make me wonder what things today are accepted that won't be in the future.
I'm not totally sure what the conclusions should be from the "ESP guy" thing. I think they include bigger questions about epistemology than simply "good research design" -- although on the other hand, research design was definitely a part of it, and pre-registration specifically actually specifically relevant to Rhine. One of the debunking claims of his critics is that he achieved his positive results by burying studies that didn't have the results he wanted (or even not being hygienic about what runs were part of what study), just keeping the ones with positive results.
Pre-registration actually might have nipped him in the bud... (unless he was willing to be just plain fraudulent of course; there are few guards against that except discovering it; I think Rhine did not think he was being fraudulent, he was just conveniently sloppy in a way he didn't think mattered... to discover fraud you have to have a common understanding of what constitutes fraud, which pre-registration is part (just part) of developing I think)
Yup. It has almost nothing to do with research quality, as Ray Hyman figured out in the 1980s. Modern ESP studies are high quality studies. It's literally only the lack of preregistration, which leads to things like publication bias and optional stopping.
This article profoundly misses the mark.
https://slate.com/health-and-science/2017/06/daryl-bem-prove...
https://slatestarcodex.com/2014/04/28/the-control-group-is-o...
It's interesting that neither of those articles on Bem mentions Rhine explicitly (at least the first one does link to him on "dating to the 1930s".)
This has all happened before. And as far as we can tell Rhine did almost exactly the same thing as Bem, cherry-picked the "successful" runs, with statistically flawed results. But Rhine was actually a lot more succesful at convincing mainstream people of his results... but while by the 1950s or 1960s nobody believed them anymore, as far as I can tell there wasn't any conclusive single turning point where everyone by consensus agreed that they had been wrong, and why. It just sort of faded away.
And isn't currently talked about much. It's sort of embaressing to scientists, I think, that so many believed in ESP from Rhine's experiments. As a result, it isn't part of general cultural knowledge, or scientific education, nor is an understanding of why Rhine wasn't right after all... and therefore, Bem is just a repeat. Or, more worrisome, the same mistakes could be made in more 'conventional' topics as well.
[I don't think either Rhine or Bem were doing intentional fraud, I think they both believed themselves. It really does look like a repeat].
People had literally been talking about preregistration for decades before the replication crisis.
That's a bit of a strong statement, but not quite untrue. There was enough legitimate interest that there was funded parapsychology research at places like SRI into the eighties that was at least serious in intent if not in execution or outcome. I just checked and the academic Journal of Parapsychology is still published, and the AAAS has gracefully not kicked the Parapyschological Association out of affiliation yet.
Yeah, the big debunking bang never happened, but the research always had a particular smell about it: smaller and smaller effects every year with greater and greater statistical significance. Preregistration helps with the optional stopping problem, but it doesn't completely solve it, and meta-analysis actively hurts. From my brief time reading the parapsychological literature and its critics for fun, I still look at any meta-analysis with extreme skepticism.
One of the results of the lack of the "big debunking bang" is that nobody seems to have learned from it, and the identical thing could happen again, as someone else pointed out in this thread, with Daryl Bem more recently, who is probably the "ESP guy" the OP actually meant (I hadn't heard of him before).
(I think both Bem and Rhine were not intentionally fraudulent, but well-intentioned... the lessons of how good intentions can lead to accidental statistical manipulation were not learned, and probably still happen all over, not just in disrepuptable areas like ESP)
[I can't seem to find it now, but I recall reading something from someone writing in the 30s, someone who is still considered a reputable scientist today, on the lines of "I didn't want to believe in esp, but I am forced to recognize it scientifically on the strength of Rhine's research." But I can't recall who!]
On the acceptance of Rhine's results,I recall seeing something similar recently, possibly here on HN, but I don't recall details either.
In Alan Turing's most famous paper, it mentioned that human ESP could be a possible loophole in the Turing test. This was how I became aware of the history of ESP research.
Turing even considered True Random Number Generator manipulation via ESP, which, as I've learned later, also had a surprisingly long history of being a standard test in this field. His response to this loophole was a funny one: If ESP is used, it may increase the correct number of guesses in both the human mind and the computer TRNG...
> I assume that the reader is familiar with the idea of extra-sensory perception, and the meaning of the four items of it, viz. telepathy, clairvoyance, precognition and psycho-kinesis. These disturbing phenomena seem to deny all our usual scientific ideas. How we should like to discredit them! Unfortunately the statistical evidence, at least for telepathy, is overwhelming. It is very difficult to rearrange one’s ideas so as to fit these new facts in. Once one has accepted them it does not seem a very big step to believe in ghosts and bogies. The idea that our bodies move simply according to the known laws of physics, together with some others not yet discovered but somewhat similar, would be one of the first to go.
> This argument is to my mind quite a strong one. One can say in reply that many scientific theories seem to remain workable in practice, in spite of clashing with E.S.P.; that in fact one can get along very nicely if one forgets about it. This is rather cold comfort, and one fears that thinking is just the kind of phenomenon where E.S.P. may be especially relevant.
> A more specific argument based on E.S.P. might run as follows: “Let us play the imitation game, using as witnesses a man who is good as a telepathic receiver, and a digital computer. The interrogator can ask such questions as ‘What suit does the card in my right hand belong to?’ The man by telepathy or clairvoyance gives the right answer 130 times out of 400 cards. The machine can only guess at random, and perhaps gets 104 right, so the interrogator makes the right identification.” There is an interesting possibility which opens here. Suppose the digital computer contains a random number generator. Then it will be natural to use this to decide what answer to give. But then the random number generator will be subject to the psycho-kinetic powers of the interrogator. Perhaps this psycho-kinesis might cause the machine to guess right more often than would be expected on a probability calculation, so that the interrogator might still be unable to make the right identification. On the other hand, he might be able to guess right without any questioning, by clairvoyance. With E.S.P. anything may happen.
> If telepathy is admitted it will be necessary to tighten our test up. The situation could be regarded as analogous to that which would occur if the interrogator were talking to himself and one of the competitors was listening with his ear to the wall. To put the competitors into a ‘telepathy- proof room’ would satisfy all requirements.
People rate him highly based on watching his spectacular, dramatic diving catches – forgetting that those dives are the product of him being badly positioned on the hit, and that a correctly-positioned player would simply make an unremarkable routine catch on the same hit.
Jeter's the baseball equivalent of the programmer who an Org is in love with for constantly spending nights and weekends fighting fires, failing to notice that it's the same person who's accidentally starting all of the fires to begin with.
Despite his -9.4 dWAR, his 96.3 oWAR still leaves him well beyond the 60 WAR HOF barrier.
Jeter is such an odd player since he is simultaneously one of the best to ever play the game and massively overrated. I’m not sure there is another baseball player who fits that description.
People underestimate how hard it is to produce offensively in a pressure cooker environment in NY. Many superstars have come here and simply not produced the way they have in the past because they can't produce under the pressure.
In a way Jeter's career is a cross of those two. His triple slash is similar to Rose, and he's also among the hits and plate appearance leaders. And like DiMaggio, he was the face of a Yankees dynasty.
The products of science are personal advancement in the field.
The products of science have monetary value to the people who fund science.
---
There are lots of reasons to "do science", but how does a particular person associated to a particular lab decide to do a particular study?
Anyway, it's hard to take any study involving humans (health, diet, economics, psychology) seriously if it's not pre-registered. I think there should be funding specifically for pre-registered studies, but I'm not king of the world. Any journal in those fields should require preregistration, but it should also be funded.
I think it's relatively common for practitioners to have an idea for some change to an existing model or system, be that a structural change, new feature, some preprocessing step, etc, develop it by iteratively making changes and running some pipeline, and stop when it looks like a measurable improvement has been achieved. Cut a PR, put in some graphs of the change in metrics of interest, merge, and start the process again. Any sequence of negative results encountered on the way are quietly presumed to be because something in the code or the setup wasn't right yet -- and often there were clear problems which get fixed along the way. But as a result, those negative outcomes often aren't kept or reexamined.
https://betanalpha.github.io/assets/case_studies/principled_...