If not: prove it. And if yes, under what conditions exactly?
If not: prove it. And if yes, under what conditions exactly?
Much better discussion of it than I offer.
abstract:
The do -calculus was developed in 1995 to facilitate the identification of causal effects in non-parametric mod- els. The completeness proofs of [ Huang and Valtorta, 2006 ] and [ Shpitser and Pearl, 2006 ] and the graphi- cal criteria of [ Tian and Shpitser, 2010 ] have laid this identification problem to rest. Recent explorations un- veil the usefulness of the do -calculus in three addi- tional areas: mediation analysis [ Pearl, 2012 ] , trans- portability [ Pearl and Bareinboim, 2011 ] and meta- synthesis. Meta-synthesis (freshly coined) is the task of fusing empirical results from several diverse stud- ies, conducted on heterogeneous populations and un- der different conditions, so as to synthesize an esti- mate of a causal relation in some target environment, potentially different from those under study. The talk surveys these results with emphasis on the challenges posed by meta-synthesis. For background material, see 〈 http://bayes.cs.ucla.edu/csl papers.html 〉
http://www.amazon.com/Causality-Reasoning-Inference-Judea-Pe...
[0]: http://www.michaelnielsen.org/ddi/if-correlation-doesnt-impl...
You need to make assumptions to be able to draw a causation from data. This is what people are doing when they design controlled experiments.
You have to assume the data you have is all that there is, and sometimes the math will only give a class of causal graphs instead of a single causal graph.
There's another issue: their approach is statistical, so even if you assume a class of graph structures you can only draw conclusions that the causal structure "almost surely" exists, not that it does exist. What this means is that even within infinite data, you can't prove that the other structures are impossible. If I turn on my sprinkler when it rains the first {very large number} times, you may conclude that there is almost surely a causal link between the rain and wet grass, but there is no reason I have to do it. You aren't left with a proof.
James Heckman [0] is one of several who make this their life's work. He has applied this analysis to fields like early childhood education.
[0] http://www.nber.org/papers/w13934
[1] http://en.wikibooks.org/wiki/Econometric_Theory/Regression_v...
Now, the challenge is that free randomness doesn't exist. But with a good approximation (i.e. it has little common cause with the dependent variable, and is caused very little by the dependent variable. a dice roll would in many cases be a very good approximation) we can get a result with high confidence.
This can be done based on data alone, if someone else has performed the experiments, or through a "natural experiment".
We can say whatever we like, really. If you mean "is it true", then yes, if our sample was representative of all the programmers and not biased in some way, the new programmer will have a 80% chance of liking science.
In reality, though, "80% chance of liking science" is just a best guess. Since all we have to go with is the sample we already studied, we assume it's representative (or try to correct for the bias we know) and make a guess based on that.
This has nothing to do with causation, however.
One way to see all the pieces is to recognize that at least in theory there are real answers to all of these questions which we are approximating. We could literally round up every "programmer" and ask them whether or not they "like science" and then get a completely accurate statistic that, say, "92% of programmers like science".
But that's infeasible (even buying that my quoted words could make literal sense) so instead we run a finite study. This is called taking the subpopulation of a superpopulation and drawing the statistic from that subpopulation. Now we see that in our subpopulation of N participant-programmers, 80% of them liked science.
Now let's bring in a new guy. Let's say we brought him in by taking every programmer in the superpopulation (replacing our sampled subpopulation back into the global pool) and picking one guy in a lottery. Provided that was a truly random lottery then we can use the 92% number from before and know that there's a 92% chance that our lottery picked a science-liker.
But we don't actually have that information. Instead, we simply have an estimate to 92%, our subpopulation statistic of 80%. If our error bars on guessing incorrectly are fine with a 12% discrepancy then we're still golden, though.
But there's one final element—we want to be able to talk about our sampling process and the risks of using it. The only parameter of this process is the size of the subpopulation chosen. For instance, if we had sampled only 5 people then there's about a 1/4 chance that we'd arrive at an estimate of 80% and a 7/10 chance we'd get 100%. This still leaves around 10% of the possible subpopulations we could choose that would tell us something really misleading, like only 20% of programmers like science.
This could be called sample statistic stability, or just straight up estimation error. It's rarely taken into account when people say things like "80% of programmers like science" because the usual assumption is that the sampled subpopulation is "big enough" to minimize this error sufficiently. In truth, though, it's hard to know exactly what the possible error of your sampling method really is.
So finally, we want to know that, given we saw 80% of our subpopulation likes science, what's the chance that a new, randomly drawn programmer also likes science. The honest truth is that there's a 92% chance, so we're ignoring some amount of error by taking our study on faith and saying that there's an 80% chance. However, given literally no other information, it's still our best bet to assume there's an 80% chance—it's better than randomly picking any other number based on the data we've seen. [0] Finally, if we were to repeat this whole experience over and over again starting with fresh, random draws of the subpopulation then we'd be, on average, correct in our guesses about the ratio of the super population [1].
[0] Bayesians here would say that we should mix in other information we might have to form a better estimate—I think that makes a lot of sense in a situation like this, especially if our subpopulation is tiny, so I'm trying to be really clear that we have no other information. Let's not talk non-informative priors.
[1] i.e. if we keep repeatedly re-estimating our "80% statistic" and using it honestly to make our guesses then it would bounce around based on our subpopulation, sometimes 80%, sometimes 100%, sometimes 20%, etc. If we multiply these estimates by how often they occur and sum it all up then we'll get exactly 92%. We'd get there faster just by drawing a bigger subpopulation, though.
The causal statement would be "training to become a programmer makes you like science more" and it already begins to indicate that we'd want to observe people who both do and do not decide to become programmers and question their preferences over time in order to have relevant information.