So the idea is to make some challenge that where there is a magician team to detect frauds, a statistics team to avoid p-hacking and minimize the chance of flukes, something equivalent to pre-registering and other security measures to ensure the result is legit.
What is the best scientifically documented EPS experiment?
As for Randi, who watches the watchers? Is he truly disinterested in his challenge, or are his examinations also biased? Additionally, the research I've seen shows a statistically significant, but faint result, so it may not be easy to reproduce in Randi's setting.
All that being said, I have not personally been able to reproduce the results, even though the claim is that almost everyone has a low level of latent ESP capability, which can be amplified through training. SRI even has an ESP training app if you want to try verifying the claims.
The first thing I though was a bad synchronization, but the anticipation was 3 seconds, that is big enough.
The second possibility is the filtering process. It's weird and it's a common source of error. The interesting part is in page 21-22:
> Low consistency responders
> Given that there was better evidence for presentiment from the consistent responders, one wonders how the inconsistent responders performed?
> Examination of the raw data revealed that in most cases, the inconsistent responders were so labeled because one or two of their calm trials had exceptionally large within-trial variances (i.e., variance of the physiological measure from the moment of the stimulus to the end of the cool-down period). Because of this observation, as a post-hoc test we examined the mean presponse correlation for SCL for each inconsistent responder, after removing the one calm trial with the most extreme within-trial variance.
> Table 9 shows that the effect of removing this single trial from each of the total of 49 inconsistent responders (across both experiments), indicated as I* in the Table, dramatically changed the overall results.7
> All data from all subjects combined resulted in a nonsignificant zpre-r = 0.03, whereas removing the single highest-variance calm trial from each of the inconsistent subjects (leaving 98.5% of all data) resulted in a combined significant zpre-r = 2.99 for SCL.
So the problem is that they are filtering the people that made a random move just before seeing a calm image, but they are not filtering the people that did a random involuntary move just before seeing an emotional picture.
Also, he has a number of such studies. Are they all suspect due to filtering, or are some filter free?
To simplify the examples, I'll assume that each person saw the same amount of clam and emotional photos, and also that half of the photos were in each class. (In the paper the number varied from person to person and the ratio was approximately 2 to 1, the conclusion is similar but it's more difficult to write.)
They use an analog sensor, but they are essentiality counting how many times a person reacted before seen an image. It can be a premonition or a sneeze or any other cause.
In the part I posted, they selected the people that reacted exactly once before a calm photo. With this selections it's not clear how many times each person would react before an emotional photo. They compare the data, and the difference was not statistically significant, so they reacted approximately once before an emotional photo.
This is somewhat a coincidence, there is no theoretical reason for this, but if you assume some sensible distribution of the chance to react randomly before a photo and use some hand waving, this is not very surprising because if people have no PSI abilities, they'd react approximately the same number of times before each set.
So now you have a bunch of people that reacted exactly once before a calm photo and approximately once before an emotional photo. We all agree that this is obviously not a proof that they have some premonition.
Then they remove the 1.5% of the reactions before a calm photo and keep all the reactions before the emotional photos.
And now you have a bunch of people that never reacted before a calm photo and reacted approximately once before an emotional photo. So there is a clear difference in the reactions before the photos and they misinterpret this as a proof of premonition.
They actually use an analog sensor, so there is more noise involved and makes everything more fuzzy. If the noise level were too high it could have overshadow the bad cut they made in the data, but the noise was not so high.
> Also, he has a number of such studies. Are they all suspect due to filtering, or are some filter free?
I don't have time to read every study he published, but if you link one (with full text) I'll try to see if I can find an error.
At any rate contra your original claim it appears a psi effect has been extensively scientifically documented.
I think Scott's analysis of "my priors against this are so strong that evidence in support of it looks more like evidence against the validity of the methodology" broadly matches with mine.
Forgive me if I am misreading you, here.
"I looked through his list of ninety studies for all the ones that were both exact replications and had been peer-reviewed (with one caveat to be mentioned later). I found only seven...Three find large positive effects, two find approximate zero effects, and two find large negative effects. Without doing any calculatin’, this seems pretty darned close to chance for me.
...
That is my best guess at what happened here – a bunch of poor-quality, peer-unreviewed studies that weren’t as exact replications as we would like to believe, all subject to mysterious experimenter effects."
Either these are some quite incompetent and/or deceitful researchers, or they have found something we currently cannot explain.
I believe that the CIA used to experiment with it, but nothing useful came out of it.