Causality in machine learning
unofficialgoogledatascience.com
unofficialgoogledatascience.com
My job was to take an inefficient proof-of-concept R package, and make it computationally and memory efficient enough to run on real-world datasets. I failed totally. The obvious explanation would be that I just wasn't able to understand the math involved well enough to implement the algorithm despite immersing myself in it for months. My personal guess though is that both the paper and the reference implementation were flawed in some way that made the task impossible.
Anyway, the paper is here:
Targeted Maximum Likelihood Estimation for Dynamic and Static Longitudinal Marginal Structural Working Models
Petersen, Schwab, Gruber, Blaser, Schomacher, and van der Laan; J Causal Inference 2015
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4405134/pdf/nih...
And the R package here: https://github.com/joshuaschwab/ltmle.
My understanding was never perfect, but my belief is that there is kernel of insight in this approach that has not yet been explored in machine learning. Alternatively, maybe it works as is, and just needs a better implementation. I'd love to see someone implement this approach, or fix it, or discredit it. As it is, I think it's potentially incredibly valuable work that is getting very little attention.
What exactly was considered wrong with the output? How did this fail?
I'm still trying to understand this, and I worry that much of it is due to my personal failings. I think I'm much worse than most people at making headway on problems when I know that my understanding is flawed, and I rationalize this as part of my moral code as a sort of "first do no harm". But I'm scared to discard my moral instincts for sake of convenience, out of fear that I wouldn't know where to stop.
In this case, I couldn't find an interpretation of the paper that matched my interpretation of the sample code, and I couldn't find an interpretation of either that seemed plausibly correct. The flaws in each seemed so obvious to me that I had to presume that I was not interpreting either correctly, and I concluded that I must be missing something essential.
There wasn't much in the way of test structure other than eyeing up the results and declaring it to be OK. I was scared that if I inadvertently implemented something incompatible with both the paper and the code, my errors would never be caught. None of the grad students seemed to fully understand the paper, and the professor wasn't familiar with the R code. My clumsiness with the terminology of the field made it difficult for me to communicate with the professor.
I feel terrible about the whole thing, but don't know what in particular I should have done differently.
I know that Petersen and van der Laan are well-respected researchers in causality that know exactly what they are doing. E.g., Petersen has a course (I think at Berkeley) that uses causal graphs which were pioneered by Judea Pearl (see comments below). I can only second the recommendation to dig into his work.
That would be exceptionally powerful if applied to the social sciences. I've long thought that the extreme lack of statistical rigor in that field has led to all kinds of damaging and biased social policies that negatively affect society, ranging from the legal system to the criminal justice system.
I'm not really denouncing the unofficial blogs, but I can see it making PR people uncomfortable, and understandably so. On a personal level, it seems like an easy way to use the name of your employer (btw, how do we know they're really employees of X) as a springboard for popularity, without the huge responsibility that comes along with officially representing the organization.
"In the following paper I wish, first, to maintain that the word "cause" is so inextricably bound up with misleading associations as to make its complete extrusion from the philosophical vocabulary desirable; secondly, to inquire what principle, if any, is employed in science in place of the supposed "law of causality" which philosophers imagine to be employed; thirdly, to exhibit certain confusions, especially in regard to teleology and determinism, which appear to me to be connected with erroneous notions as to causality." http://www.hist-analytic.com/Russellcause.pdf
http://bayes.cs.ucla.edu/BOOK-2K/causality2-epilogue.pdf
>Take, for instance, Newton’s law:
>f = ma.
>The rules of algebra permit us to write this law in a wild variety of syntactic forms, all meaning the same thing – that if we know any two of the three quantities, the third is determined.
>Yet, in ordinary discourse we say that force causes acceleration – not that accelera- tion causes force, and we feel very strongly about this distinction.
>Likewise, we say that the ratio f/a helps us determine the mass, not that it causes the mass.
>Such distinctions are not supported by the equations of physics, and this leads us to ask whether the whole causal vocabulary is purely metaphysical, “surviving, like the monarchy . . .”.
>Fortunately, very few physicists paid atten- tion to Russell’s enigma. They continued to write equations in the office and talk cause–effect in the cafeteria; with astonishing success they smashed the atom, invented the transistor and the laser.
It just seems like discussion of "what causes what" is missing the point, just like the above argument. That one persisted for centuries (millenia?) until it was pointed out the answer is neither! It depends on your reference frame.
Lets return to your equation. r^2 = m1*m2/F is a statement of correlation that's valid to many digits, that also says that "these are all the variables that affect anything here." It allows for calculations of quantities by plugging in numbers. In combination with other equations, it allows extrapolitaion and physical thinking. But equations alone are not enough for a physicist to make progress on problems, they don't combine arbitrarily, they still use that famous "physical intuition" that turns out to involve a heck of a lot of causality that they talk about in the cafeteria but don't often write about.
In the world outside of physics classrooms, most situations involve a huge number of variables interacting with each other, and the variables themselves are leaky abstractions. This means that the correlation of quantities does not go out to many digits, or even one digit. This does not make an equation a useful abstraction. People have developed probabilistic models that stand in for equations in these situations. Sure, it would nice to have an equation that modeled things like they do in the simple situations that Physics publishes on, but when the equations of physics are so weak that they can't even fold a protein, how are we going to use them to model the interaction of a potential drug with a protein?
Equations are great, when they apply. By introducing simplifying assumptions, we can sometimes extend their utility into new domains, but at the cost of introducing fragility because of those simplifying assumptions. They allow for calculating new things with only a very few computations. But nearly every new field of science invents new math. One of the more recent inventions is causality. It works fantastic in those sciences that use experimental interventions. Simple algebraic equations do not.
Physicists might pray together in the cafeteria to come up with good ideas. It doesn't mean God answered their prayers when they have one.
Useful laws need not be exactly correct. You can keep adjusting to stay on course to account for the error due to any simplifications.
Sure causality seems like it must be at least some kind of heuristic, that doesn't mean looking for causes is the best, or even a good, use of time.
Fair enough.
>"I spend a lot of time in the world of correlations, in biology."
I left that field. One reason was because I realized so much of my time was wasted after I attempted to attach numbers to things like "# of receptors required to fit the narrative".
Sure "A causes B", but it must be for some reason totally different from what is usually proposed (and could probably be considered a boring experimental artifact). I just don't care if "A causes B". I want to know for "f(A) ~ B", what is f?
>"It's much better to get a causation: a knob that when you turn it you get a response, rather than no response."
Reminds me of this paper:
"Biologists summarize their
results with the help of all-too-well recognizable diagrams,
in which a favorite protein is placed in the middle
and connected to everything else with two-way arrows.
Even if a diagram makes overall sense (Fig. 3a), it is usually
useless for a quantitative analysis, which limits its
predictive or investigative value to a very narrow range.
The language used by biologists for verbal communications
is not better and is not unlike that used by stock
market analysts. Both are vague (e.g., a balance between
pro- and anti-apoptotic bcl-2 proteins appears to control
the cell viability, and seems to correlate in long-term with
the ability to form tumors) and avoid clear predictions."
https://bml.bioe.uic.edu/BML/Stuff/Stuff_files/biologist%20f...If you try to model f(A) ~ B in biological systems without simultaneously modeling causality, then it's going to go wrong. This happened to me in year 2 of grad school. The bayesian network built around data without modeling the causality was wrong, the one that took into account the causality from the knockout experiments was right.
I wonder if you'll recommend your doctor use that bit of wisdom the next time you're sick.
It would actually be miraculous to me if it is the former, to the point I would be tempted to believe there is some occult driving force leading humans in the right direction.
However, I was thinking about similar the other day regarding the dentist. Say I go get a teeth cleaning. Did the polishing cause my teeth to be cleaner, or did the cleaner teeth cause the dentist to do the polishing? I feel like there is some kind of semantic game being played there.
Wow...
Eh, maybe it is a dumb example, but I'd appreciate if you could put your finger on it rather than sarcasm.
Minor nit: the equations aren't weak at all — they predict the gyromagnetic moment of an isolated electron to ten decimal digits of agreement with experiment. The weakness is in the computational implementation of those equations, i.e., solving the equations in a reasonable time. You can perform a relativistic quantum monte carlo simulation of protein folding, and I guarantee the results will match those obtained in the lab to whatever level of accuracy you desire, but you'll have to wait longer than the age of the universe for the computation to finish running.
Also, for work on small proteins, the calculated chemical properties match the experimentally tested ones very closely nowadays. The problem is all of the unknown interactions that a new drug has with the rest of the human body outside of one protein. And without a large atomistic model of a human body, it's going to be a long time before simulation can solve that problem.
This is precisely what I mean by "weak." The equation looks great on paper, but isn't so useful in practice. And the problem isn't predicting just one protein-drug binding, it's predicting the interaction with all known proteins, all molecular complexes, and the biofilms, etc. I would go so far to say that Physics with a capital P is not so useful for many physical problems at all. Where Physics is useful, it's unbelievably useful, but its domain in the physical world is vanishingly tiny.
Physics kind of has the reverse problem of AI. With AI, when you finally solve a problem, it's no longer considered AI, and AI is continued to be considered a failure. With Physics, a problem isn't part of the body of Physics until it's been solved, at which Physics is now considered a great success.
Someone is new to the idea of relativity. This is nonsense. The frame of reference describing the sun rotating around the earth is not inertial (https://en.wikipedia.org/wiki/Inertial_frame_of_reference) and thus does not follow the laws of physics we know and love. In short, "Nope, you cannot say the sun orbits the earth."
Is that not the case?
d = sqrt(dx^2 + dy^2 + dz^2)
F = G*m1*m2/d^3
That would just be summed up over all the objects included in the system. However, when visualizing the results, I could set (0,0,0) wherever I wanted. I could put it at the solar system barycenter, I could put it at the center of the sun, I could put at at the center of the galaxy (this leads to interesting vortex-like trails...), or I could put it at the center of the earth. The easiest was actually to use the solar system barycenter on an arbitrary date.To be clear: Pearl is not just philosophizing. He provides a formalization that takes causality from being an ambiguous concept to being applied mathematics. The bulk of the book is not a light read, but the mathematics involved are relatively basic.
And to be absolutely clear: Pearl won the Turing prize for his work. It is not simply handwaving or non sequiturs. The common online conversations about causality are several decades out of date with real academic work on the topic.
Regarding the Turing prize, etc. That type of thing doesn't really convince me. What engineering feats or surprising predictions has Pearl's methodology lead to? When I look at what is successful in machine learning these days, it is RNNs, CNNS, gradient boosting, etc. Correlation-based methods seem to be king, at least for now.
Pearl's work is heavily used: https://scholar.google.com/scholar?cites=1099626011922949961...
The Russel quote you posted is in no way some sort of rebuttal of Pearl et all's work. They aren't even in the same ballpark. I don't intended that as a slight agains Russel who I'm generally fond of.
> Recent progress in artificial intelligence (AI) has renewed interest in building systems that learn and think like people. Many advances have come from using deep neural networks trained end-to-end in tasks such as object recognition, video games, and board games, achieving performance that equals or even beats humans in some respects. Despite their biological inspiration and performance achievements, these systems differ from human intelligence in crucial ways.
> We review progress in cognitive science suggesting that truly human-like learning and thinking machines will have to reach beyond current engineering trends in both what they learn, and how they learn it. Specifically, we argue that these machines should (a) build causal models of the world that support explanation and understanding, rather than merely solving pattern recognition problems; [...]
https://arxiv.org/pdf/1604.00289.pdf
Some others papers related to these ideas:
- Learning a Theory of Causality http://stanford.edu/~ngoodman/papers/LTBC_psychreview_final....
- Theory-Based Causal Induction: http://cocosci.berkeley.edu/tom/papers/tbci.pdf -
From the brief summary it sounds very hypothetical, especially compared to the causality-agnostic pattern-recognition approach which I have seen work surprisingly well. That group has convinced me they are on to something quite big.
For example in this review: https://arxiv.org/pdf/1605.01138.pdf, a model-based approach is compared against a deep learning approach. It is very interesting to me that even if deep learning performs better, its failure cases are very different than human failure cases (deep learning is not sensitive to "physics illusions"). Whereas the model-based approach is very much in-line with human performance characteristics. Also, on this task (physical scene understanding), deep learning generalises very poorly.
That is an interesting connection to make though. If people view wild speculation as the alternative focus of research, incorporating the concept of causality would at least slow them down. This would reduce the amount of misinformation generated and number of destructive policies implemented. Still, I think causality offers more than that.
You appear to be saying that causality is not quantitative, however that is not true.
Namely, in that a rational belief is justified on the basis of prior experiences, while a rationalization is justified on the basis of impracticality due to ignorance.
Here is a spielraum[1] of possible experimental results, with the * indicating the region consistent with a law:
1) |------*------|
Here it is for a vague speculation that there is "some correlation" (ie astrology): 2) |******-******|
You can see it will be much easier to find evidence consistent with #2 than #1, even if there is nothing to the ideas at all.[1] Paul Meehl. 1990. Appraising and Amending Theories: The Strategy of Lakatosian Defense and Two Principles That Warrant It. Psychological Inquiry 1990, Vol. 1, No. 2, 108-141. http://rhowell.ba.ttu.edu/meehl1.pdf [figure 3]
But it has a weakness - it doesn't always understand causality. As you note, much of the time this doesn't matter.
A good counterexample is the story of scurvy prevention[2]. Back in 1747, the British Royal Navy proved that scurvy could be prevented by eating citrus. The Navy started serving all sailors lime juice, and sure enough scurvy stopped occurring.
Fast forward to 1911. Scientists had developed rules which matched the observed behavior of scurvy very well:
Atkinson inclined to Almroth Wright’s theory that scurvy is due to an acid intoxication of the blood caused by bacteria.. There was little scurvy in Nelson’s days; but the reason is not clear, since, according to modern research, lime-juice only helps to prevent it. We had, at Cape Evans, a salt of sodium to be used to alkalize the blood as an experiment, if necessity arose. Darkness, cold, and hard work are in Atkinson’s opinion important causes of scurvy.
Basically, what had happened was that widespread steam travel and a process for canning lime juice were developed roughly simultaneously. The canned lime juice no longer provided vitamin C, but no one realized this was a problem because sea voyages were much shorter so scurvy never developed.
That lack of a correct causality model meant that when polar exploration happened, scurvy started occurring again and no one knew why.
[1] https://www.mapr.com/blog/association-rule-mining-not-your-t...
But my point was to show that causality is important, which the OP appears to doubt altogether.
A non-causal regression model on {Y={CA, no CA}, X ={Blonde hair, dark hair}} may be able to replicate the observed effect that blonde hair and skin cancer go together. However you cannot use that regression coefficient for coming up with an intervention. Dyeing the hair is not going to reduce cancer risk although the regression model will predict exactly that.
http://ftp.cs.ucla.edu/pub/stat_ser/r350.pdf
"If correlation doesn’t imply causation, then what does?", Michael Nielsen
http://www.michaelnielsen.org/ddi/if-correlation-doesnt-impl...
Experiment. Prove of implementation, the way molecular biologists do it.
P(Dog barks | do(kick the dog) )
The `do` modifies your model to treat "kick the dog" as observed, breaking its dependence on any other variables, e.g. the doorbell ringing or the cat hissing. With those links broken, the model becomes simpler, and often the causal question can be answered using observational rather than experimental quantities.
At least, that's the idea. Not sure how broadly accepted it is, but Pearl is a towering figure and his work seemed robust at least to my not-so-expert eyes.
There are some principles which cannot be undone by any amount of hipsterism and sophisticated sectarian bullshitting. There is no way to jump from observation to causality without knowing the implementation. At least in this particular universe.
In the realm of models (or ideas) it could be seem doable, but a map is not a territory, model does not represent reality until proven experimentally.
For example, in the social sciences it's common to look for "instrument variables" which are random and can't be affected by anything else. If you can find these types of variables then you can pretend that you're looking at the results from an experiment.
There are also methods like discontinuity analysis, which argue that if your system is discontinuously effected by a variable, but the causes will effect the variable in a continuous way, then you can look at changes around that specific point as being causal.
I don't see why you need to design a new experiment if data that satisfy your requirements already exist.
For example, if I wanted to know whether people given the name 'george' are more or less likely than the population as a whole to have at least one child, must I recruit pregnant women, and assign half of them to name their child 'george' and allow the others to choose any name? Would it not be better to look at existing data to compare Georges and non-Georges, perhaps slicing by socioeconomic variables? The results would be instant (rather than requiring 65 years to gather) and the dataset larger.
A is correlated to C. B is correlated to C. A is NOT correlated to B.
How is that possible?
The argument is that, if causation is unidirectional and acyclical, then the only causal structure that leads to the above is that A and B both cause C.
There's no other way to do it. And you can derive causation from observational data!
Of course in real life there are a bunch of problems with this, particularly measurement error and hidden unmeasured or unmeasurable variables. But more or less those exist with experimental data as well. You can only conclude things about observed variables and a very large number of unmeasured variables could throw off your inference.
Anyway, I more or less held your belief on this until I read Pearl's (and colleagues') material. There's much more expansion to it than the canonical example above.
Churches are related to people. People are related to the Sun. The Sun is caused by the prayers of the people in churches. Sun is in the sky today because someone somewhere prayed.
At the extreme end of the spectrum, you have to weigh the probability that the statistical evidence is high enough to suggest causation against the probability that your mind isn't plugged into a simulation and all of the measurements have been faked. If the former is greater than the latter, than I'd say for all practical purposes we can call the phenomenon "causality".
Hidden context, uncontrolled experimental variables, pleotropic effects, experimental noise all conspire to produce conclusions that cannot be reproduced or if reproduced, cannot be generalised. The contrast with physics is stark - theoretical physicists predict experimental results (eg Higgs boson) many years before the experiment is performed. I don't think that in biology, we merely need to prove a succession of hypotheses to achieve the kind of explanatory causal models we want. A theoretical framework is needed, and at the moment, there is not even an outline of such a framework.
It's possible to detect the direction of causality with observation with some additional assumptions that are reasonable general.
For example Additive Noise Models (ANM) assume that there is an additive noise structure in observational distribution. The key assumption is that if X causes Y, the noise in X can have an effect on Y but not vice versa.
Additive noise model is near 80 per cent accurate in correctly determining cause-and-effect across large number of datasets.
---
Distinguishing Cause from Effect Using Observational Data: Methods and Benchmarks Joris M. Mooij, Jonas Peters, Dominik Janzing, Jakob Zscheischler, Bernhard Schölkopf; 17(32):1−102, 2016. http://jmlr.org/papers/v17/14-518.html
Center for Causal Discovery web site http://www.ccd.pitt.edu/
Fine. What if you missed a hidden variable Z which also causes Y? What if there is also X1 which causes Y when no X is present.
For example, a mapping of genetic mutations to actual diseases cannot be done from purely observational data without knowing the implementation - how particular proteins interact in this or that pipeline.
Possible causal relationships does not establish or prove causality itself.
Noise going trough Z -> Y is not present in X -> Y.
Detecting X1 when X is not present is trivial. Y is not present and the noise from Y is not present. There is another cause besides Y.
>Possible causal relationships does not establish or prove causality itself.
Noise models can establish the arrow direction with very high probability. The strength of causality and other factors are considered separately