http://ftp.cs.ucla.edu/pub/stat_ser/r350.pdf
"If correlation doesn’t imply causation, then what does?", Michael Nielsen
http://www.michaelnielsen.org/ddi/if-correlation-doesnt-impl...
http://ftp.cs.ucla.edu/pub/stat_ser/r350.pdf
"If correlation doesn’t imply causation, then what does?", Michael Nielsen
http://www.michaelnielsen.org/ddi/if-correlation-doesnt-impl...
It's possible to detect the direction of causality with observation with some additional assumptions that are reasonable general.
For example Additive Noise Models (ANM) assume that there is an additive noise structure in observational distribution. The key assumption is that if X causes Y, the noise in X can have an effect on Y but not vice versa.
Additive noise model is near 80 per cent accurate in correctly determining cause-and-effect across large number of datasets.
---
Distinguishing Cause from Effect Using Observational Data: Methods and Benchmarks Joris M. Mooij, Jonas Peters, Dominik Janzing, Jakob Zscheischler, Bernhard Schölkopf; 17(32):1−102, 2016. http://jmlr.org/papers/v17/14-518.html
Center for Causal Discovery web site http://www.ccd.pitt.edu/
Fine. What if you missed a hidden variable Z which also causes Y? What if there is also X1 which causes Y when no X is present.
For example, a mapping of genetic mutations to actual diseases cannot be done from purely observational data without knowing the implementation - how particular proteins interact in this or that pipeline.
Possible causal relationships does not establish or prove causality itself.
Noise going trough Z -> Y is not present in X -> Y.
Detecting X1 when X is not present is trivial. Y is not present and the noise from Y is not present. There is another cause besides Y.
>Possible causal relationships does not establish or prove causality itself.
Noise models can establish the arrow direction with very high probability. The strength of causality and other factors are considered separately
Experiment. Prove of implementation, the way molecular biologists do it.
P(Dog barks | do(kick the dog) )
The `do` modifies your model to treat "kick the dog" as observed, breaking its dependence on any other variables, e.g. the doorbell ringing or the cat hissing. With those links broken, the model becomes simpler, and often the causal question can be answered using observational rather than experimental quantities.
At least, that's the idea. Not sure how broadly accepted it is, but Pearl is a towering figure and his work seemed robust at least to my not-so-expert eyes.
There are some principles which cannot be undone by any amount of hipsterism and sophisticated sectarian bullshitting. There is no way to jump from observation to causality without knowing the implementation. At least in this particular universe.
In the realm of models (or ideas) it could be seem doable, but a map is not a territory, model does not represent reality until proven experimentally.
For example, in the social sciences it's common to look for "instrument variables" which are random and can't be affected by anything else. If you can find these types of variables then you can pretend that you're looking at the results from an experiment.
There are also methods like discontinuity analysis, which argue that if your system is discontinuously effected by a variable, but the causes will effect the variable in a continuous way, then you can look at changes around that specific point as being causal.
I don't see why you need to design a new experiment if data that satisfy your requirements already exist.
For example, if I wanted to know whether people given the name 'george' are more or less likely than the population as a whole to have at least one child, must I recruit pregnant women, and assign half of them to name their child 'george' and allow the others to choose any name? Would it not be better to look at existing data to compare Georges and non-Georges, perhaps slicing by socioeconomic variables? The results would be instant (rather than requiring 65 years to gather) and the dataset larger.
A is correlated to C. B is correlated to C. A is NOT correlated to B.
How is that possible?
The argument is that, if causation is unidirectional and acyclical, then the only causal structure that leads to the above is that A and B both cause C.
There's no other way to do it. And you can derive causation from observational data!
Of course in real life there are a bunch of problems with this, particularly measurement error and hidden unmeasured or unmeasurable variables. But more or less those exist with experimental data as well. You can only conclude things about observed variables and a very large number of unmeasured variables could throw off your inference.
Anyway, I more or less held your belief on this until I read Pearl's (and colleagues') material. There's much more expansion to it than the canonical example above.
Churches are related to people. People are related to the Sun. The Sun is caused by the prayers of the people in churches. Sun is in the sky today because someone somewhere prayed.
At the extreme end of the spectrum, you have to weigh the probability that the statistical evidence is high enough to suggest causation against the probability that your mind isn't plugged into a simulation and all of the measurements have been faked. If the former is greater than the latter, than I'd say for all practical purposes we can call the phenomenon "causality".
Hidden context, uncontrolled experimental variables, pleotropic effects, experimental noise all conspire to produce conclusions that cannot be reproduced or if reproduced, cannot be generalised. The contrast with physics is stark - theoretical physicists predict experimental results (eg Higgs boson) many years before the experiment is performed. I don't think that in biology, we merely need to prove a succession of hypotheses to achieve the kind of explanatory causal models we want. A theoretical framework is needed, and at the moment, there is not even an outline of such a framework.