When correlation is better than causation
narrator.ai
narrator.ai
> [Abductive reasoning] starts with an observation or set of observations and then seeks the simplest and most likely conclusion from the observations. This process, unlike deductive reasoning, yields a plausible conclusion but does not positively verify it. Abductive conclusions are thus qualified as having a remnant of uncertainty or doubt, which is expressed in retreat terms such as "best available" or "most likely". One can understand abductive reasoning as inference to the best explanation.
> In the 1990s, as computing power grew, the fields of law, computer science, and artificial intelligence research spurred renewed interest in the subject of abduction.
Abductive reasoning is basically how one would formally describe 1) the practice of medicine, including diagnosis, 2) the rules for evidence in legal trials, 3) the process for generating hypotheses in science, and innumerable similar activities we undertake daily.
And for obvious reasons there's a close relationship between abductive reasoning and Bayesian statistics.
"Inference to the best explanation" could mean we accept any explanation regardless of how improbable it is - as long as it best explains the data.
The bayesian idea is that we can learn something about causation if we accept uncertainty and impose "sanity constraints" (priors) on the explanation.
Without knowing the real-world mechanics of Y, we can say something like "setting X to 0.33 will increase Y, with 60% probability." It maybe impossible to learn anything else from the data.
Malcolm Gladwell's similar message: https://www.pushkin.fm/episode/burden-of-proof/
- A correlation between mining and lung cancer was discovered in 1918, but wasn't acted on until 1975.
- There is a correlation between football and suicide/brain damage, but it is not being acted on.
https://www.espn.com/nfl/story/_/id/22603654/nfl-doctor-says...
https://www.today.com/parents/brett-favre-psa-urges-no-tackl...
Whether the game can ever be made safe is another issue.
It's profoundly sad that institutions of higher learning are promoting activities that they know can cause brain damage and long term disability, just so they can make money and entertain their alumni.
This raises a broader point about Collinearity and whether correlation is actually actionable when the feedback cycle is long. You could easily be working the problem for 20 years before you ever knew you were wrong.
https://mattasher.com/2020/04/29/the-filter-podcast-episode-...
Personally, I'm not quite positive that I buy that causation is that hard to establish in many cases. Don't give up on that idea. One thing I would say is that if you have a strong prior reason to believe that one thing causes another thing, finding that they are strongly correlated, that can be a useful datum. Mainly the important thing is to understand the limitations of correlation to guide decisionmaking.
Does that make sense? Smaller steps to make sure you only invest if it is worth it
[1]: https://www.microsoft.com/en-us/research/wp-content/uploads/...
To understand Minka's expectation propagation algorithm you might first need to get a little intuition about assumed density filtering. One way to understand assumed density filtering could be to read a few tutorials about hidden Markov models [3] or Kalman filters and try to get a feel for why and when and how people might want to approximate posterior probability distributions. It might be hard to build enough intuition without trying to actually apply the things (implement the algorithms) or prove the theory yourself, and then try to come up with your own ideas for how to improve the algorithms.
I completely agree that there are basically infinitely more things to learn than available lifetime. It helps a lot to have a concrete application or goal in mind: then you can focus on learning the tools and theory that move you closer to the goal, rather than learning bits and pieces of unrelated knowledge that don't connect together in a useful way.
[1] https://tminka.github.io/papers/ep/roadmap.html
[2] Minka's EP slide deck from his PhD defense https://tminka.github.io/papers/ep/defense.pdf
[3] Rabiner wrote a famous HMM tutorial https://courses.physics.illinois.edu/ece417/fa2017/rabiner89...
I read Causality which I understand is the more technical of their books, and everything was presented surprisingly intuitively. Sure, I had to go over some things twice, but that's to be expected when you learn something new.
If you're worried, start with one of the more pop-aimed books? You'll be fine.
(Pearl did change the way I look at causality and correlation, fundamentally for the better, so I do strongly recommend getting familiar with it. I also liked Willful Ignorance which is sort of one the same theme but also not and takes a wider approach.
A lightweight introduction to Pearl's ideas is the epilog of his book, which is also his Turing award lecture. Here's a pdf scan, there's also video of him giving this lecture up on the internet if you prefer: http://bayes.cs.ucla.edu/BOOK-2K/causality2-epilogue.pdf
And therein, I believe, lies the problem.
I think the issue is the pressure for science to produce something constantly so in today's world, correlation is causality. Whether or not you believe in deterministic laws that govern reality, correlation is often the easiest approach when looking at a difficult problem and there in lies the rise of much of probabilistic and statistical models in the face of difficulty. Not all cases, but a lot of cases. We don't want to continue trying the hard work of determining definitive casual relations, if they exist and are content with correlative relations.
As someone who grew up fascinated by science because it was science that sought and provided causal relations, I'm often disappointed about the current world of research. I'm not saying this work is easy by any means, it just seems like we often give up anymore after we pick up the low hanging fruit.
I'll use one example from some data I've been looking at, which is whether the covid-19 pandemic has changed how people sleep. To study this using the formal notion of causality requires asking a random 50% of people to sleep as if covid isn't happening. That's obviously both impractical and implausible.
So you can really only look at correlations. But I can show you the correlations, and I bet you will be convinced that the pandemic HAS changed peoples' sleep. Here's some charts if you can take a look: https://jeffhuang.com/covid_sleep/ but there's probably several factors that convince you that this is causal.
First is the pattern of sleep pre-covid is very stable, and feels trustworthy because it goes up and down during weekends, and holidays are visible. So the data is visibly sensitive to changes in the environment. Second, nearly every country reacts similarly when the N is separated, so even if there's some large group of people somewhere that are outliers (say, some policy by California that everyone needs to go to bed later), it would only affect that one country they are in, not each country separately the same way. Finally, the patterns of sleep post-covid are also stable with similar patterns as pre-covid, but just shifted.
I'm not sure if there's formal ways of representing these concepts, but I feel humans understand these intuitively.
https://www.pnas.org/content/117/24/13386
It's similar in a business context - the "pressure" of finding a causal result (especially in situations using AB testing) lead to poor analysis practices in order to find something significant.
Working in data, especially in a startup, we often need to make so many decisions and trying to change the culture is good but when it’s a fire then this approach would get us the furthest
Do you mean it's not hard from an analytical point of view / from a practical data gathering perspective - or both?
I've found that needing causation leads to big delays in backlogs, especially when it's required for every insight, but I'm curious if you've seen it to be different.
1. Abductive: https://en.wikipedia.org/wiki/Abductive_reasoning ex: Hypothesis
2. Inductive: https://en.wikipedia.org/wiki/Inductive_reasoning ex: Generalizations
3. Deductive: https://en.wikipedia.org/wiki/Deductive_reasoning ex: Studies
You don't prove causation, but you can disprove it when you find absence of correlation.
Observed correlation suggests causation which allows you to make a prediction. A prediction can be tested. The prediction will either be true or false based upon whether the correlation continues to hold.
This is one of the problems with A/B tests--they often don't have causation aka "Why?" "This dialog box was rearranged and gave us 15% better conversion." Um. Okay. But "Why?" If you can't answer "Why?" you don't have causation.
"We removed needing to enter a phone number and now have 15% better conversion." "Why?" is obvious in that case.
You're correct that they can't tell you the the root cause of why your change causes a particular difference, but that's a separate issue from correlation and causation.
Perfect is the enemy of good as they say.