The Easiest Data Analysis Mistake to Make
blog.statwing.com
blog.statwing.com
It's also worth mentioning "dumbbell" data sets. Two clusters of data, each of which have a independent, meaningful correlation in them, can easily leverage a linear regression into a meaningless line passing through the two clusters. That's a pretty common issue with high dimensional data (obviously you can see it in a 2D scatter plot), and it's not easily caught short of looking at regression diagnostic statistics.
As you point out obviously you can see it in a 2D scatter plot, but you have to select the correct two variables.
Right now I do my data analysis in numpy, but this looks good for my Excel-based colleagues.
What library is doing the statistics?
EDIT: clarified libraries we use.