The Case for Causal AI
ssir.org
ssir.org
For example, can I answer questions like what are the main causes that are moving corona stats?
With intervention I suppose you can conduct an experiment where you (randomly) pick cities and make half of them wear masks and half of them not. But this is of course unethical! And some stuff you want to know the effects of (e.g., how did gatherings at protests affect the infection rate) are one time events that can't be replicated again.
So all we have left are lots of natural experiments. Different countries/states/communities are handling the situation differently with a wide range of outcomes. No two communities are directly comparable since they differ along many other dimensions other than corona policy. But as a human I am still drawing plenty of conclusions on what caused what and to what degree. So it seems like a solvable problem. How do I make it rigorous?
Pearl's "Book of Why" may be a good starting point as a popularization, although sometimes Pearl is not so easy to read. And Pearl presents his own perspective on causal inference only -- there are other schools and techniques, although there are generally equivalences between them.
It is definitely possible to understand the core ideas as a layman, if you have a mathematical grasp of basic probability. If you want to go a little further, understanding linear regression helps too.
The problem is that actually applying the theory to draw causal conclusions on real-world data seems to be a very subtle and difficult process and even experts who specialize their entire careers in causal inference frequently make mistakes and disagree with each other. So I think a lot of humility is warranted.
So, the solution is to get comfortable with the lack of rigor. What can be known in a system without rigor? That’s the question to make rigorous, I think.
Who is working on making that rigorous?
The problem I have with causality is how to find accurate causal diagrams from unstructured observational data. We _could_ guess and check every possible causal relationship, but that’s at least exponentially hard—in which case causality is useless.
People seem to have reasonably good capabilities for generating candidate causal hypotheses from observational data (basically all of modern science), but most of the material I’ve found on causality focuses on the theoretical benefits it provides rather than on practical applications at scale. (I don’t care if we can automate finding the causal graph for barometer & air pressure; how do I find a causal graph for classifying fake/real news from plain text data?)
Control theory models handle these things just fine. But control models are hard to apply to sociological/epidemiological domains, where causal inference dominates.
From what I gather, causal inference is useful for designing studies. I’m not sure if they’re used for prediction — would appreciate if someone in the know could chime in.
I don't think this should be true, and if it is, then "causal inference" should be qualified to refer only to a specific modeling framework. As a counter example, it's possible to formulate a nonlinear differential equation model of Covid spread and infer the parameters to construct a plausible, causal, generative model.
> There are two approaches to causal AI that are based on long-known principles: the potential outcomes framework and causal graph models. Both approaches make it possible to test the effects of a potential intervention using real-world data. What makes them AI are the powerful underlying algorithms used to reveal the causal patterns in large data sets. But they differ in the number of potential causes that they can test for.
Does anyone have references and tutorials for either approach?
The program is the explanation of the output, i.e. the program causes the output.