Graphics That Seem Clear Can Easily Be Misread
scientificamerican.com
scientificamerican.com
The graph in [1] explains it much better than any words could.
[0] https://en.wikipedia.org/wiki/Simpson%27s_paradox [1] https://www.analyticsindiamag.com/understanding-simpsons-par...
No relation to The Simpsons; it is Simpson's paradox, after Edward Simpson (e.g., https://ftp.cs.ucla.edu/pub/stat_ser/r414.pdf).
Not sure about the ad placement in that analyticsindiamag.com article.
Aside: it's not always true that the direction of correlation in the sliced groups is more correct than that of the larger group. You need causal analysis to know when it is incorrect or correct to slice the data to eliminate confounders.
Perhaps you mean that drawing any kind of fit though this data is misleading, when lacking confidence statistics and R2 score - although you could have all those and still do poorly on held-out data, or you could even do well on held-out data and still be victim to confounders or misunderstanding the direction of causation...
I see a line through data as nothing more than a hypothesis, which needs to be backed up with rigor. If it's backed up, there's nothing outrageous about it.
One example in which the combined perspective is the valid one is election results (it is possible to win a higher percentage of votes in multiple areas, yet lose the overall vote).
Simpson's Paradox can strike even the most committed truth-seeker trying to understand and interpret the available data.
Edit to add: Simpson's Paradox could conceivably thwart a gerrymanderer if they have an erroneous model relating demographics to voting patterns.
I hadn’t thought about it that way or seen the word gerrymandering used in that way, but it makes sense.
Its to do with websites. Lets say there a website where one downloads some free software from, for example. If there's a really big bright clear download link, Ive been taught to not click on it, as its probably going to an ad or some scam site. Ive learned to always ignore it and search for the tiny little text hyperlink that actually leads to the real download file.
Ive noticed this trained behavior of mine now applies to all websites, even when the big link is actually the real one. I will often miss it, and waste time search for the 'real' one as my brain now automatically filters the big clear ones out.
A typical example is "given any city, there is a strong correlation between number of churches and number of crimes commited." This is pretty much true everywhere in the world but that does not imply one is causing the other. This correlation can simply be the natural outcome of more populous cities having both in higher numbers compared to less populated ones.
While that form of thinking may sound incredibly stupid, example: how could a person confuse a cause for its effect, it is exceedingly common. I have seen incredibly smart people make this mistake. The mistake is the non-cognitive behavior at play that unduly influences what is otherwise a very logical and straight forward conclusion. Objectivity is a practicable personality trait not aligned to logic or math skills.
Like, I know what "correlation" is: slap a regression on A and B and see what comes out (after considering heteroscedasticity and friends). But, for causation, how does one find it?
Is causation even a well defined concept? In most disaster analysis situations you see that failure wasn't caused by a single factor but it came as a result of a combination of different factors which, on their own, are benign.
If someone is crossing the road while checking thir phone (so not paying attention) and a drunk driver hits them with their car, what "caused" the accident?
Do phones "causes" accidents? Does drunk driving "cause" accidents?
What's more, the real take-away is that you can put side by side two graphs. It doesn't mean there is any causal relationships between them. The example only seems convincing because both graphs are health-related. If they were Pac-man highscore vs milk production, it would show how hollow the article is.
Alberto Cairo is quite the polarizing figure in the data visualization world.
"How to Lie With Statistics"
https://www.amazon.com/How-Lie-Statistics-Darrell-Huff/dp/03...
The page now loads correctly
In this case, I haven't read the article the chart was taken from; I don't know what argument the chart is supposed to support. Stripped of that context, it's a pretty confusing chart. It seems the author (of the SciAm article - Cairo) is using the chart to make a point about lying with charts. I don't think that Cairo is publishing his research on obesity and birth-weight - that seems to be Kitahara et al. If that's what's going on, then it's hardly surprising that the chart is hard to read; Cairo chose it to make exactly that point.
And I think he's being unfair to Kitahara et al., implying that they've deliberately contrived that chart to mislead.