Could It Be? Spooky Experiments That 'See' The Future
npr.org
npr.org
I would be a lot more relieved if people with the power over other peoples lives grokked Baye's rule and priors.
Bayes's rule doesn't help with the point that suggestive evidence is not convincing evidence. It just points out that prior beliefs are part of the equation, but will hopefully pale in comparison to actual data. In fact, I was taught to set practically useless hyperparameters to ensure that they do. No one does that outside of an experiment.
Let's say I believe (I don't) the height of pygmies is normally distributed, where the mean is also normally distributed with mean 130cm and standard deviation 10cm, and the standard deviation is inverse gamma distributed with shape 7cm and scale 1cm. Assuming the height is actually normally distributed with mean 160cm and sd 15cm (it isn't), how many pygmies must I measure to admit that P(height>160cm)>20%? I'm not sure I can even do the math.
Here P=50% for the unknowable accurate model and P=0.13% for the prior model. How does the situation change when my prior is "sufficiently close to 0"?
In the legal case example, maybe some clarity maybe had in considering what does the prior mean. A answer is: say you have to bet a million dollars on whether the person is guilty or not without knowing anything about the person how would you distribute your million dollars between the two events. Yes it is subjective and personal, but it is hardly ever going to be 50:50. One can push the $1,000,000 analogy further. One can fix a cost for a mistake: whats the cost of a wrong conviction and whats the cost for setting a guilty man free. Then the final decision can be based on reducing the financial risk based on the likelihoods.
One may bring the socio-economic status in forming the priors but one may not consider any information source that considers the accused.
Lets take the example. There is a one to one correspondence with fictitious counts and priors. One way of encoding a 50:50 prior is to construct a possibly fictitious but representative past of (say) 2000 samples split into 1000 guilty and a 1000 not-guilty. After each prediction and assuming that the truth gets known one has to update the counts appropriately, so that the next time we use a different prior.
Our initial prior may be wrong but it will approach the correct one asymptotically. But how fast it approaches the true prior depends on how wrong our initial prior was.
It makes the very good point that this paper makes lots of statistical tests, and then bases big claims on the small minority that showed a significant effect. This is in no way restricted to psychology, drugs companies for example do this all the time. It's cheating, whether you realize it or not.
Statistical significance only tells you that a result is unlikely to be a fluke -- not that it definitely isn't a fluke -- but the more tests you do, the sooner you'll see a fluke on average.
In other words, if you toss a coin 1,000 times, then it's hideously unlikely that you'll see a run of 100 consecutive heads. But if you toss the coin 100,000,000 times, you shouldn't be too surprised to see that 100-toss run buried in there somewhere, even though the odds of getting 100 in a row are so small.
Right?
The only way to find out is to do enough flips to eliminate the chances of your final result being influenced by statistical flukes. Measuring small differences, like trying to answer "does a coin preferentially land on one side vs the other?" usually takes hundreds of thousands of tests to guarantee you're seeing objective data, rather than seeing patterns in noise.
the very moment you peek, your data is tainted from future testing.
If you toss 100,000,000 == 2 ^ 27, you should only expect around 30 in a row. To have a good chance of getting 100 in a row, you need about a billion squared times more.
And the problem is MUCH worse than described above: Let's say you test 1000 wrong hypothesises with p=0.05; 50 of those will be accepted as true, even though all are wrong. If you test 980 wrong hypothesises and 20 right ones, more than half of those that pass the p=0.05 "golden" significance test will in fact be wrong.
Now, when you see a medical journal with 20 articles using p=0.05, which do you think is more probable - that 19 are right and one is wrong, or 19 are wrong and one is right? The latter has a much higher likelihood.
The whole field of systematic reviews and meta-analyses has developed around the need to aggregate results from multiple studies of the same disease or treatment, because you can't just trust one isolated result -- it's probably wrong.
http://en.wikipedia.org/wiki/Meta-analyses
Statisticians working in EBM have developed techniques for detecting the 'file-drawer problem' of unpublished negative studies, and correcting for multiple tests (data-dredging). Other fields have a lot to learn...
Regardless of the true reason, these are never carried out before a new drug or treatment is approved (because there is usually one or two studies supporting said treatment, both positive).
And if you have pointers to techniques developed for/by EBM practitioners, I would be grateful. Being a bayesian guy myself and having spent some time reading Lancet, NEMJ and BMJ papers, I'm so far unimpressed, to say the least.
It's when the hypothesis predicts a pattern that you haven't noticed yet and the pattern is confirmed by experiment, that's when you know you have something [a theory].
But to be fair, we should take things on their merits. Contrasting the two papers, Bem's paper sticks very closely to the standard form and 'rules of engagement' (if you will) of published academic research. This counter paper on the other hand, has only the surface appearance of proper scientific method. For example, in making one if its main points it refers to a hypothetical casino where Bem could've made an 'infinite' amount of money, and to the one million dollar Randi prize, neither of which reach the standard of proper experimentalist scientific evidence. It does this by way of justification to one of their main points which is that because 'extraordinary claims require extraordinary evidence' we should be able to set the prior expectation of H1 - in their words 'for illustrative purposes' - to .00000000000000000001, which they then go on to demonstrate, makes the results non-significant.
But where does the 0.00000000000000000001 come from? It could, just as easily, be twice or half that figure. Thus not falsifiable and therefore not justifiable as an extra bit of arithmetic that Bem's paper must qualify.
To put this in terms this audience will understand, that's kind of like saying well no, because you are using Java and 'everybody knows' Java is slow I think we should multiply your benchmark figures by, oh, let's say, one half. And then we see that java throughput is quite poor, as expected. In fact, we should multiply all java benchmarks by some number like a half but I'm not going to be specific about it because actually its pretty much just an arbitrary number I made up. So, reading past the reference to Bayes and some nice formulas that's just arithmetic in my book.
Note, I'm not saying that there are no flaws in the Bem paper. Everybody can see that it's very likely they'll be something wrong with it (and the desk drawer problem to which the above paper also makes extensive evidence is a likely though not conclusive contender) but I think its only reasonable to hold up the criticism's to the same standard as what they are criticizing. Perhaps that way you'll be more likely to find the actual truth of the matter.
That's kind of the problem: published academic research is broken in a lot of ways. If you haven't read it yet, the New Yorker's recent article on the decline effect touches on several of the reasons why: http://www.newyorker.com/reporting/2010/12/13/101213fa_fact_... And some prominent medical journals started fighting against publication bias a few years ago: http://www.smh.com.au/articles/2004/09/09/1094530773888.html
I think you're holding this response to a much higher burden of proof than it needs to meet to be a proper refutation of Bem's claims. You're right that the response doesn't appear to use the "proper scientific method". But that's because it doesn't use it at all, and it doesn't need to. There's no hypothesis to test and no experiment to run in order to point out flaws in a paper that does claim to be the result of the scientific method.
Unfortunatley, publishing these kind of claims prematurely help the more gullible among us to fall for ridiculous claims from psychics and others who would take advantage of them. (The authors of "The Secret", I'm looking at you.)
Commenters on NPR's website (not exactly the dumbest audience online) have already shown this problem; "All of you criticizing this need to open up your minds" and "The future, as well as the past, influence our dreams."
This type of attitude is infuriating - if a claim can survive the crucible of peer review then it we can be much, much more certain that it is true and correct. If humans were to posses a limited form of precognition that would be awesome - but before we can claim to posses a thing we must be sure that it is real.
Disease existed long before we figured out causes and targeted treatments; people possessed it, and people applied treatments with varying degrees of success and failure, including death. One of the more interesting inaccuracies of medical history is that of humors (random site - http://www.gallowglass.org/jadwiga/herbs/WomenMed.html). How we treated disease changed over time and is still changing as we learn new things.
My point? It is important to continue to apply rigid scientific study to all manners of phenomena, not only to validate its existence but also to figure out how to repeat or avoid said phenomena, depending on the need, the positives and the negatives of said phenomena. However, we should not turn a blind eye towards what people think they experience just because we have not yet come up with the right tool for measuring or the right study for identifying. There is always some reason behind the claim (even if the reason is "snake oil salesman").
Whatever is behind precognition (to take your example), people claim to experience it and always have made those claims. There is a certain burden of proof required, sure. How do you convince someone born deaf that there is a thing like sound that is experienced the way the hearing experience it? In the case of precognition, it tends to be self-validating (sometimes self-fulfilling), and yet it is still useful for precogs or those who believe in them, whether it is illusion or real, whether we have proven it concretely or not.
That bears a quick repeat... The information is somehow useful. These people who are shouting "open your mind" find their precognitive information useful; in their minds, challenges to this useful information are silly. In the name of understanding, the real focus should be on figuring out how that information is obtained. Is it psychic phenomena, a ghost whispering in the ear, great subconscious brain processing, or something else?
So if the response to an outright simple rejection is "open your minds", I think it is warranted. If the response is to indicate disagreement, however, I always thought that to be a useless response, as useless as the simple rejection.
True. But on the other hand, publishing ridiculous claims and incorrect results is a necessary part of science.
When we publish only results we know to be correct, because they agree with mainstream beliefs, we introduce a bias into the scientific process. In reality, if you publish 20 experiments with p=0.05 [1], 1 of them should be incorrect. If less than 1 in 20 of your papers isn't wrong (assuming p=0.05 is the gold standard), you are not doing science.
You can see a perfect illustration of this when people tried to reproduce Millikan's oil drop experiment. I'll quote Feynman: Millikan measured the charge on an electron...got an answer which we now know not to be quite right...It's interesting to look at the history of measurements of the charge of an electron, after Millikan. If you plot them as a function of time, you find that one is a little bit bigger than Millikan's, and the next one's a little bit bigger than that, and the next one's a little bit bigger than that, until finally they settle down to a number which is higher.
Why didn't they discover the new number was higher right away? It's a thing that scientists are ashamed of - this history - because it's apparent that people did things like this: When they got a number that was too high above Millikan's, they thought something must be wrong - and they would look for and find a reason why something might be wrong. When they got a number close to Millikan's value they didn't look so hard. And so they eliminated the numbers that were too far off, and did other things like that...
This is why I'm an advocate of accepting/rejecting scientific papers based solely on methodology, with referees being given no information about the conclusions and with authors being forbidden from post-hoc tweaks. You do your experiment, and if you disagree with Millikan/conclude that ESP exists, so be it. Everyone is allowed to be wrong 5% of the time.
[1] I'm wearing my frequentist hat for the purposes of this post. Even if you are a Bayesian, you should still publish, however.
People occasionally mention p-value calibration and note, sadly, the damage caused by this reckless practice that allows false results to eke through the airtight, 300' tall walls of scientific publication. But there is value in being wrong. It's a part of science.
In a way, it's the MD's White Coat syndrome applied to PhDs. Something that is scientific and written in a journal is necessarily correct in public opinion instead of the rigorously considered opinion it really is. Both paper-reading public and the authors of some of those papers tend to believe this.
And to cover it from a Bayesian point of view, it's pretty vital to keep the culture such that the risk of publishing something incorrect doesn't to strongly dominate the decision to publish. You should be confident talking about your beliefs long before they distribute like deltas.
This would be simple to do by submitting ethics applications and an analysis plan to a trusted third party which would only release them once the author is in a position to publish, or at a pre-agreed cutoff (perhaps 2 years), whichever is the shorter (to avoid scooping). Perhaps I should set something up...
And I suspect the ultimate reason it's not done this way... is that scientists in certain fields would publish a lot fewer papers, not slightly fewer but a lot fewer, if all the effects they were studying had to be real.
When some people find that their model doesn't quite fit, they make a more accurate model. Others make a less specific model. It's the difference between model parametrization and model selection.
So when we get a dubious result, we can either say "no result" or "possible result". The choice tends to depend on how the finding affects future research. Biology is more exploratory than confirmatory, so they go that way.
Instead, I should have focused on science journalists, who should be extra diligent when reporting these sorts of stories to point out the possibility that this is a false positive.
A million times yes. Also: no publication without the experiment's methodology and criteria for success having been registered prior to the experiment's commencement.
In reality it doesn't turn out this way because the results that get written and published tend to be biased in favour of novelty and demonstrating a relationship rather than the absence of one. How many similar experiments could have been terminated, never submitted or not published because they failed to show anything notable? This is one of reasons... 'Why Most Published Research Findings Are False' http://www.ncbi.nlm.nih.gov/pmc/articles/PMC1182327/
That meta-study applied to medical studies and I think this genre would probably fair even worse when it came to long term replicability.
I agree, but we need to be really clear about what the claims are.
It could be that there is a reproducible 1% "mystery" effect that works from future to past, but only in experiments like this. In which case claim wouldn't be extraordinary, it'd just be something we can't understand.
Remember that he's still in the data gathering stage. If -- and it's a big if -- there is any kind of reproducible pattern that doesn't match known laws, that doesn't mean there is a claim. There is simply data that doesn't fit our current models.
That's why people who push frontiers have to be very, very careful about differentiating the data from the claims.
Lots of guys make lots of money with bogus TV shows and books on stuff like this, and it's a shame: many times there is something unusual in the data, but the claims jump far ahead of any reality. Fear of this effect has probably silenced a lot of little tiny pieces of data that wouldn't make sense -- it's simply too much trouble to have to keep explaining yourself. This might be one of the factors explaining Feynman's story of the Millikan oil drop experiment.
Erm, no, that would be pretty damn extraordinary. We know of nothing else in the universe that acts like this.
Good scientists know there is much yet to be discovered.
The idea that the effect could be the result of a programming error or a small amount of light leaking through/around the screen is completely plausible. Or it could be dumb luck as you suggest, but it's extraordinarily unlikely.
Specs: Two buttons - ESP mode or non-ESP mode. In non-ESP mode, 60 random pictures are shown and we are to guess. Then it gives the correct one. In ESP mode, add some porn. If the results are different, then we have ESP. Use Javascript for the randomisation algorithm so that we can be sure there is no server trickery being done.
Eg, I once had a Poker playing game on an Amstrad that used a PRNG with a very predicable seeding strategy. I could amaze my friends by knowing exactly what cards I would be dealt.
>I think you need a true random number generator to truly test this
EDIT: or http://www.lavarnd.org/
As such, test subjects knowing they are in the group with the erotic pictures should still be able to beat the 50% threshold.
Edit: Actually, I'm reading through the actual experiment and it appears all 100 sessions used both erotic and nonerotic pictures of varying arousal value. Also, both the position of the picture and the picture itself were not actually chosen by the computer until after the test taker made the choice, although they were told differently, making it a test for a future event. So, yes, I think we would already be compromised for trying to recreate the test.
In other news, capital punishment has been installed for science journalists publishing articles that contain a question in the title that can be succinctly answered with "No."
edit: just to note, nowhere in constructing a statistical test is it required that the creator decide how "extraordinary" the null hypothesis is.
To be persuaded of something you already strongly believe is far easier than something you don't believe. And really the key word is persuade. It's not that people can prove the sun will come up tomorrow, but they can persuade you that it will. .
If you think that's illogical, I'd ask you to consider why a teacher is more likely to accept the excuse "my dog ate my homework" than "aliens kidnapped me and stole it". You seem to be arguing that given that the evidences are equal (a mere statement from a kid), the teacher should properly consider both occurrences to be equally likely.
In your example, based on existing data, it is indeed fair - some dogs do sometimes eat homework, whereas there are no verified accounts of aliens stealing it. So that's a legitimate adjustment of priors. Particularly if you actually have data on the incidence of paper-hungry dogs.
But in science and philosophy, there's lots of important questions for which we can't legitimately calculate priors, and "it would be too weird" is not at all relevant when determining their values.
The line "extraordinary claims require extraordinary evidence" is just more poetic than the paragraph above.
Now, let's say ten different scientists are interested in this claim, and they're all going to run their own experiments. The chance that all ten will run an experiment with each reporting "false" is under 60%.[0] Over 40% of the time, at least one scientist will falsely conclude the existence of the phenomenon that definitely does not exist. This is an effect of running multiple independently-considered experiments without aggregating the results.
That's the Bayesian problem that people mention. Another problem entirely comes from which results will tend to get published.
Now let's consider the effect of publishing bias. Let's assume that only 20% of the scientists will attempt to publish their results regardless of the outcome, but they will always try to publish if the (false) phenomenon is shown to exist. This effect alone results in 21% of submissions being incorrect,[1] even though an incorrect result only has 5% likelihood.
Let's additionally assume that a journal will publish a false-but-interesting result 50% of the time, and the true-but-ho-hum result only 10% of the time. The final effect is that 50% of published results for this extraordinary-but-false phenomenon incorrectly report the phenomenon to be true.
Tweak the numbers all you want, but the effects of running multiple independently-considered trials, along with biased publishing, means that we are surprisingly likely to publish false conclusions.
Notes:
[0] 0.95^10 = 0.5987
[1] 0.95 * 0.2 = 0.19; 0.05 / (0.19 + 0.05) = 0.21
[2] (0.21 * 0.5) / (0.21 * 0.5 + 0.79 * 0.1) = 0.50
53% in an experiment that has 36 trials? Really? 50% is 18/36, but 19/36 is 52.7%, and 19/35 is 54.3%. Depending on experimental design, if it just so happens that near the end of the 20 minutes, the subject has a tendency stop on one of the last erotic pictures they're like to guess and let the time run out.
Hypothesis #2: Depending on how the computer's random number generator was seeded (and they might have a relatively short repeating sequence), subjects may have, however unconsciously, "learned" to predict the randomness, something they would have insufficient motivation to do in the other set of pictures. [We can test for this by seeing if they were getting better at it over the course of a session.]
Highly recommended if you're thinking about making a podcast.
http://www.npr.org/blogs/krulwich/2010/12/08/131910930/neil-...
http://en.wikipedia.org/wiki/Funnel_plot
PS non-physicists invoking "quantum stuff" are bullshit merchants, p < 0.001 :-)
My wife (a psychologist and a Christian) defended the (mostly psychologically oriented) experiments as posed in the paper and the method behind them, whereas several atheist/strictly-causality-believing coworkers and also a conservative Christian with a strong anti-psychology bias dismissed the idea of spooky action from the future outright. The stormy e-mail exchange raged on (and I did not contribute to the discussion). The conservative guy accused me of abandoning my wife in the argument. I believe she's well capable of handling herself.
In any case I wrote this (bad) poem in response:
My act is mostly mute and unseen,
I wear no costume, I don’t vent my spleen.
Spending most time behind the stage,
Conceiving a plot I prepare the cage.
I have few resources, can’t sponsor M-M-A,
Must find some other way to while away the day.
I step out briefly to address the crowd,
I hope today they’ll surely be wowed.
My mind’s been active, reading Hacker News,
What’s this I see? Some interesting views,
on whether the future can affect our present,
I’m sure this will stoke plenty of dissent.
I have my materials for a good time today,
Setting the stage is just an e-mail away.
My fingers fly fast, the idea’s not hokey,
My actors will soon be addressing the spooky.
I press ‘Send’ and my time on stage is done,
I’ve set the parameters, now it’s time for fun.
The actors appear to have done my bidding,
I just hope it doesn’t end in too much bleeding.
Sure, I’ll show up from time to time,
The audience gets tired of hearing everyone whine.
They need to see larger schemes at play,
The actor’s philosophies won’t save the day.
Arguments, screeds, reasoning galore,
It’s exciting for a time, not yet a bore.
I’ll step back just about now,
It’s time for some more.
This audience of one will now sit back,
Got a few more bugzillas to whack.
I won’t make it to peer-reviewed journals,
But empirically it’s great to see what sprouts from a kernel.From the article : "The sequencing of the pictures on these trials was randomly determined by a randomizing algorithm … and their left/right target positions were determined by an Araneus Alea I hardware-based random number generator."
At the very least they were using Araneus Alea wich is a hardware random number generator, so the numbers were not predictable. It's possible that the "randomizing algorithm" did something dumb and made the sequence not random, but I doubt it.
I think that it's more likely thet the sudy was done so many times that it eventually gave significant results than it is that the sequence was not random. Or maybe prescience is real to some degree, or the study is a statistical glitch.
However the replication package they provide has the compiled program without the source code, and that is a red flag to me.
1. Buy ticket
2. Go AWOL until after the drawing.
3. Study numbers
4. Look up results
5. Profit
http://norvig.com/experiment-design.html
here on HN again. How well does Bem's set of experiments hold up?
http://psychsciencenotes.blogspot.com/2010/11/brief-note-dar...
"The real lesson? This is the level of methodological scrutiny every paper should receive, and not just the ones you think are crazy: the ones you like and rely on for your own work should get a good working over like this too (especially these ones; and I'm as guilty on this as everyone else)."
A (future) message from <your favorite porn provider name here>
Here's one IEEE paper: "a perceptual channel for information transfer over kilometer distances - historical perspective and recent research" http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=1454...
Can't find the online version anymore, so here's the version I found (9.5MB pdf): http://www.box.net/shared/inxg0nld9r
1. People recalled as many as possible.
2. People were divided into two groups, control and experimental.
3. Experimental group was asked to retype the words.
4. Control group was asked to go home.
5. Experimental group found to score better on initial test.
I do not know if that makes a difference, but I think it might and could. The retyped words might have whatever quality or for whatever reason might have been easier to rememmber than the non retyped words.
Unless in the original article it suggests that the experiement was carried out in the way you suggest, I do not think there was much control of variables.