A Philosopher Reviews Judea Pearl's “The Book of Why”
bostonreview.net
bostonreview.net
Separating out causality opens the door to doing counterfactual reasoning in better ways. Since whatever method is in the Book of Why confused the reviewer, it's probably not the best way to look at counterfactuals, so the causality bits can be taken without establishing the counterfactual methods as the be-all end-all.
For that matter, DAGs versus cyclic directed graphs or time-based DAGS is still a big concern too. Many if not most of our causal reasoning will have loops, when time is not accounted for, and DAG formulations make this difficult to unroll. There may be big improvements in causal modeling, without considering the counterfactuals even.
Also, in Pearl's prior 2000 book on Causality, it's clear that Pearl gave plenty of credit to Spirtes, and it always seemed that they were working in parallel and on very similar problems; I'm not sure how much Spirtes took from Pearl, but Pearl makes clear that his ideas are heavily informed by Spirtes' work.
"This happened because I did that" seems to imply "Had I not done that, this would not have happened" -- a counterfactual.
Couldn't you have a scenario where your action was the direct cause of an outcome that might have come about some other way without your involvement at all?
For instance, if a ball rolls down a hill because you pushed it, but a breeze blows a moment after you pushed the ball, if that breeze was strong enough to push the ball down the hill too, then you have a scenario where "that ball went down the hill because I pushed it" is true, but "had I not pushed the ball, it would not have rolled down this hill" is not true.
> "This happened because I did that" seems to imply "Had I not done that, this would not have happened" -- a counterfactual.
The point is that constructing the counterfactual requires a more than just knowledge of a single causal link, but a broader set of causal knowledge about many causal links.
Language is messy and ambiguous and we do often colloquially use "X caused Y" to imply the truth of the counter factual "If X had not happened Y would not have happened". Interestingly, a different tense such as "X causes Y" generally does not have the same counter factual implication.
The point isn't just about causality. As an undergraduate, one of my philosophy professors (Bill Lycan) told a class "never attempt to define anything in terms of counterfactuals. No matter what it is, there will be decisive counterexamples."
To be super-clear, this is not a statement about Pearl, and it's not denying that counterfactuals are interesting, just a claim that you can't define much of anything in terms of them.
P.S. He actually said "analyze", but the way philosophers use that term, it's appropriate to make the substitution to avoid confusion.
Perhaps we should call Pearl's Rung 3 counterfactuals "subjunctive" or something entirely else to distinguish the concepts and methods of Rung 2 and Rung 3. But these rungs are extremely different in formulations and models, and are clearly distinct. If you want to call rung 2 "simple counterfactuals," well I'm not super into that sort of debate as long as we all agree to the meaning of the terms. And if philosophy finds the distinction between Pearl's Rung 2 and Rung 3 difficult to accept, then it may also mean that philosophy has not yet discovered what Pearl has formulated.
I think counter factual reasoning or thought experiments are fundamental to the way we humans think about the world.
https://statmodeling.stat.columbia.edu/2019/01/08/book-pearl...
The review had some good call outs but given the above quote and a few surrounding criticisms, this review is more pretentious than anything.
But RCT fails -- gives a nonsense result -- when the hypothesis under test is incoherent. This wouldn't matter, except that RCT is routinely used in such circumstances, and the results treated as gospel by people in positions of authority.
Consider: Outcome X may have six causes A-F. RCT tests B, and finds that varying B only affects one in six cases. With infinitely many trials, the relationship resolves, but with one trial the difference is indistinguishable from noise.
Substitute a medical symptom for X, and a medical treatment that addresses one of six causes, for B. After one RCT, B is "shown" to be ineffective.
The problem is not B. The problem is that X is ill-defined. X could be a mental illness, or tumors in a given organ. How often do we read "anti-depressants shown ineffective"? The only way we have to distinguish one variety of depression from the next is which treatment works.
The problem is not limited to medicine.
Your beef isn't with RCTs, or with the notion of RCT as evidence. Your beef is with researchers who don't know how to think.
Your complaint is like saying, we should stop considering hammers to be the gold-standard of nail-hitters because some idiots use them to try to turn screws.
Not really. It's just enough that we don't tell the red car drivers that they're part of a special group. Just let them think we have equal number of different colors assigned to different drivers (and don't let them see what the others got). Then, the fact that their assigned color happens to be red will hold no significance to them related to the test then...
They, the drivers, shouldn't know anything about the experimental design in the first place. Let alone equal number of different colors.
https://samharris.org/podcasts/164-cause-effect/
#164 - Cause & Effect - A Conversation with Judea Pearl
August 5, 2019
In this episode of the Making Sense podcast, Sam Harris speaks with Judea Pearl about his work on the mathematics of causality and artificial intelligence. They discuss how science has generally failed to understand causation, different levels of causal inference, counterfactuals, the foundations of knowledge, the nature of possibility, the illusion of free will, artificial intelligence, the nature of consciousness, and other topics.
Judea Pearl is a computer scientist and philosopher, known for his work in AI and the development of Bayesian networks, as well as his theory of causal and counterfactual inference. He is a professor of computer science and statistics and director of the Cognitive Systems Laboratory at UCLA. In 2011, he was awarded with the Turing Award, the highest distinction in computer science. He is the author of The Book of Why: The New Science of Cause and Effect (coauthored with Dana Mackenzie) among other titles.
Twitter: @yudapearl
I first learned of Judea Pearl by stumbling across the transcript of a talk he gave while I was researching DAGs: http://singapore.cs.ucla.edu/LECTURE/lecture_sec1.htm . The way he grounded his talk in the history of thought hooked me, and the talk serves as a good general overview for those deciding if they want to pick up the book.
It point out the cliche, "correlation is not equal to causation."
It is a real thing but throwing that quote around without reading into the research or having proper foundation to assess a conducted research is just as bad.
However, the account of traditional statistics given in this article is misleading. Randomisation is a major part of traditional stats, and it is inherently a causal hypothesis: breaking the links between unobserved covariates and treatment regimes.
An important alternate contemporary causal inference framework by Rubin has origins in a 1923 thesis...
...but the content of Pearl's approach seems superior; if you ignore the academic spats.
Yes, randomization is central to classical statistics, but no, it is not inherently causal. Drawing a random sample from a bivariate distribution (X,Y) is key to doing a lot (though not all) of classical statistical inference (think of estimating slopes in regression), but the randomization does not imply anything about the causal relationship between X and Y. When you speak of randomization in the context of "treatment regimes," you are thinking about randomized controlled trials, which the piece does analyze explicitly, in some detail. So in this sense the account given in the essay is not misleading.
Anyway, the section you're pointing to agrees with me. It just happens to be overlooked when they summarise...
Say you want to do basic linear regression: you want to estimate the slope for Y regressed on X. The most stringent form of inference works like this: you draw a random sample (X_i, Y_i), modeled as n independent and identically distributed realizations from the joint distribution (X,Y). Etc. This is certainly a stochastic model; we require randomization (or some approximation of it) to do inference. But it has nothing whatsoever to do with causality.
Worth finishing?