Understanding Causality Is the Next Challenge for Machine Learning
spectrum.ieee.org
spectrum.ieee.org
https://www.hsph.harvard.edu/miguel-hernan/causal-inference-...
Don't get tangled up in the philosophy of causality. That's not the immediate problem.
> I would really love all biology students to read Elliott Sober's "Apportioning Causal Responsibility", Susan Oyama's "Causal Democracy and Causal Contributions in Developmental Systems Theory", and Richard Lewontin's "The Analysis of Variance and the Analysis of Causes".
https://twitter.com/hanemaung/status/1321488068717719552?s=2...
These are available through research gate for those without library access.
The trouble is, outside of these 1st-order effect dominant, deterministic environments, causality becomes much harder. In complex systems, stochasticity, nonlinearity, feedback loops and higher-order effects dominate. There's also emergent behavior -- properties that are true in the small are not true in the large.
Consider a complex system like human society -- can we truly determine causality of broad interventions? Likely not in a first-order way like in the physical sciences. We can do it imperfectly through tools like causal inference (Rubin) which makes much more modest claims about the "strength" of effects (average causal effect). Randomized Controlled Tests (RCT) is another tool for making causal claims.
But in a complex world, 2nd, 3rd and higher order effects dominate and so the notion of root causes itself becomes ambiguous. Richard I. Cook once said "post-accident attribution to a 'root cause' is fundamentally wrong". Though humans are attracted to the idea of a chain of simple causes (which is why we have the myth of Mrs O'Leary's cow kick over a lantern and starting the Great Chicago Fire of 1871), there's typically no easily-identified root cause. First-order causal thinking assumes a Directed-Acyclic-Graph (DAG) idea of a causality chain which converge into a set of effects, but the reality is that such a DAG, if it can be represented, is likely to be infinitely complex in a complex environment. First-order causal thinking is an insufficient mental model in a complex environments.
Instead, I think instead of aiming for a deep understanding of epistemic causality (where we try to know and represent causality), it's probably more useful to focus on instrumental causality (where we aim to know the main points of leverage that are effective in changing the system). I think we'll likely get very far just by finding the knobs that have the most effect on the variables we would like to change (that don't also simultaneously change variables that we wouldn't want to change).
[1] to determine causality, we typically have to perturb the system -- determining causality through observational data is possible, e.g. via natural experiments, but there are many epistemic restrictions which limit the claims that can be made.
Couldn't agree more! For computer models to understand causality, it must be able to interact with the environment and probe it. I think understanding causality is one and the same as reinforcement learning, where a computer model learns to interact with its environment
The Michelson-Morley experiment was enough data to get special relativity, for example.
As you suggest, the treatment effect and econometrics literature currently is on a semi-parametric trend: Given that they don't believe that one can actually produce a believable complete causal model (or DAG), one tries to estimate a treatment effect that does not depend on parametric or functional assumptions.
It has been a truism to some that "most efforts in policy are responses to previous efforts in policy." In SPC you can measure noise and overcontrol. With a long-lived policy-making institution [unusual?] honesty might admit to facing effects of previous well-intentioned policies.
Plasma physics & fusion energy (my field) is challenging for exactly these reasons (despite being 95% classical physics). It's very rare that we can do a nice controlled experiment where only a single variable is changed at a time. I joke that it's really a subset of biology, not physics.
> probably more useful to focus on instrumental causality
I partly agree -- we humans do seem to get by on rudimentary reasoning. On the other hand, the issue of back-progagation is quite similar to identifying 'root causes.' There's also an issue of combinatorial explosion of the number of possible sets of variables that interact with each other, coupled with the fact that data becomes exponentially sparse as the dimension of the space grows. The human ability to detect causal relations is really stunning when you realize how tough it is -- I wouldn't want to bet that we can reproduce it by trial and error. Evolution had plenty of time to get it right, but we don't.
Yes, identifying root-causes from the outcomes is a model inversion problem (here's the data, find the generating model), and model inversion problems tend to be ill-conditioned. In complex systems, there's so many possible combinations of causes that could have led to a single outcome that characterizing the entire set is difficult.
One should mention the anthropic principle in this context. This would then lead one to speculate about a connection between observing causal order and self-awareness.
I suspect that you made that comment because you think in the terms of rigorous detection of causality, not everyday effective detection of causality heuristically.
All wetware based self driving systems on the road use heuristics in their wetware system.
You can't hope to have self driving without heuristics. That's what deep learning is.
I don't think our fallibility in causal reasoning makes it useless to pursue as a goal in artificially intelligent systems. It doesn't need to be perfect, just useful and better than not having it. Afterall, our perception systems are pretty fallible too, otherwise things like optical illusions wouldn't exist.
Somewhat tangentially, I recall a psych paper where the researchers found that people perceive images in mirrors to be located behind the surface of the mirror. The researchers apparently thought that the image was at the surface of the mirror (ie, they misunderstood basic optics), and concluded that they'd discovered an optical illusion.
I take this to mean that we have a notion of "effectiveness", and that consequences are attributed to preceding effective actions.
Reinforcement learning is uniquely positioned to build machines that understand cause-and-effect on their own because the algorithm is allowed to interact with the world, observe the results, gather more data, rule out hypotheses, and so on.
Though I agree that this is an important next challenge (and causality has been the "next challenge" for at least the last 5 years), it's often more easily solved these days with mixing in human expert knowledge to the equation (that is, using ML alongside of human expertise).
It seems that any theorem that rules out learning causality from observational data alone would also rule out learning causality from any kind of interactions.
Unless you're assuming that agent A "knows" it has free will so its own actions have no cause, while agent B can't tell whether the environment caused agent A's actions or vice-versa. But if that's what the proof hinges on, it's pretty shallow, because agent A has no such guarantee that its own choices have no root cause.
What I meant is that you cannot learn from "general" observational data, unless it is structured in a certain way (the randomized controlled trial mentioned in a sibling). RL is able to gather data on its own, while other ML methods must do with what they are given. This means that RL could eventually discover the causal relationships, while non-RL cannot (except if the data comes from a RCT).
Because in statistics, causal inference is certainly possible without RCTs.
1. First let's get through the easy part: reinforcement learning (RL) is not unique in its ability to identify cause-and-effect - this was achieved long ago through the use of randomized controlled trials. RL merely streamlines the task of reacting to such information (as well as optimizing experiments w.r.t. a desired goal).
2. Now the trickier part: you can learn causality from observational data alone if you combine this with understanding of a mechanism. Indeed, the whole field of causal inference is an attempt to formalize and extend such methods.
There is a massive space of problems where experimentation (whether old-fashioned A/B testing or more sophisticated online learning) is simply not possible, whether because of ethical reasons, cost, or other reasons behind non-destructive study. These problems are common in medicine, economics, physics. In such problems the only data is observational. Causal inference is very valuable here.
Let me try again: the most interesting and frequent setting is where you cannot control the experimental conditions and you have no idea about the causal mechanism at play. And that is where traditional ML cannot help, but RL can.
It should not be underestimated that our collective knowledge of causal inference with statistical methods is good, but still improving.
I mean, in that sense, ML is helpful. Heck, it is already used in causal inference in two-step estimators and the like.
Or better yet (stretching and sidestepping how you meant it), give your friendly neighborhood statistician/econometrician just a dataset and they can't do causal inference. Give them propensities and column names/descriptions and a writeup of the experiment/where the data came from, and suddenly they might be able to do causal inference. It points to a need to augment our observations with more structured metadata, if we want to do causal inference with data that's just lying around.
As others have pointed out, this is kind of the point of much of Pearl's work on causality. Specifically, do-calculus provides a set of primitive operations that can be used to convert queries in interventional/counterfactual (causal) distributions to estimands in a purely observational distribution.