Torture your data long enough and it'll tell you anything
businessweek.com
businessweek.com
When I was conducting scientific research, the goal was to come up with an air-tight (or as air-tight as possible) case for your hypothesis. If you presented your findings at a meeting, you better be prepared for the onslaught of questions like "Did you consider X?" and "What about Y?".
Then I moved to the business side and holy crap are the standards lower. Of course it's easier to prove something in a lab than in the real world, but in so many cases I've seen somebody say "If you do X, you will get Y result, based on the data I analyzed". Then I raise my hand and say "But what about Z? That could explain your results." and all I get is blank stares like I just solved a differential wave function in my head.
It's a great topic, especially for businesses and startups. I think the problem is much, much deeper than correlation != causation. The basic problem is that we don't understand how to deal with statistics, especially aggregate numbers. This is a funny way to make a point, but the problem is waaaay deeper than just confusion about correlation. The errors in scientific studies, for instance, are just one example of the harm caused by these kinds of cognitive blind spots. (I say blind spot instead of lack of math education because I don't believe the root problem is a lack of understanding math. In my opinion, something else is going on.)
Sometimes, knowing what caused something is the necessary answer, and for those, a root cause analysis and proper experimental design for validation are important. But sometimes, in business and in life, just knowing that things hang together can be pretty handy.
Correlations are important clues. The entire "recommendation" world, from Amazon's collaborative filtering to Hunch's "everything you might be interested in" are all predicated on correlations.
No argument, saying correlation implies causation is bad. But it's just as bad to say "therefore, correlation is bad". DanielBMarkham's article and this BW.com post both show that it comes down to interpretation of what the data says. It's understanding the limitations of what a number, or a trend, or even a distribution can reveal. It's understanding what regression to the mean actually means, or why we consider a distribution "normal"... and that outliers actually can be profitable.
And it's a recognition that with the democratization of big data, it will get worse before it gets better... but it will get better. 40 years ago, no one ever saw the stock market on the news, or had access to it's ups and downs every second. We now all have a better understanding of stocks (well, ok, that's a bit of a stretch, but you get my drift), and their dangers. Similarly, as we get used to seeing lots more data, and discovering that if you interpret it wrong, bad things happen... well, I expect more folks to ask that next level of questions. Not all, and not much past that... but it will be a start.
I'd go so far as to say the problem is 100x more complicated than "Correlation != Causation". Given a set of factual statistics it's not terribly difficult to present them in a truthful, reasonable manner than support any side of a given argument.
So the damned lie of statistics is pretty subtle, you just have to omit the number of variables you actually looked at when you present your data.
Stats are open to interpretation, which is why academia favors peer review, where faulty underlying assumptions can be checked.
This is silly. All probability distributions are cadlag, so how can you even teach probability without the notion of right continous with left limits, which means you have to resort to limits & derivatives => Calc.
Actually, the argument for combining Calc & Stats is very compelling, because there is too much synergy. How can you teach a continous probability distribution like say the Gaussian without teaching how to integrate under the curve for the cumulative distribution function, or obtaing the probability density function via the derivative, or obtaining the variance aka second central moment via the moment generating function, which means you now have to teach atleast some fourier transforms which again means Calculus. At both UChicago & Stanford where I learnt all of my probability, calculus was quite intertwined with the teaching of probability. I believe its the same case in most other schools as well.
Without calc in probability, you can do "lame" stuff like discrete distributions ( Binomial, Poisson etc....but even there, the key insight is to show how the CDFs of the discrete distributions, which will generally have terribly complicated formulae with giant factorial expressions, can be very nicely approximated by the continous distributions for large n, small p etc. ( aka continous correction http://en.wikipedia.org/wiki/Continuity_correction ). So for a large number of coin flips trials, you use a Normal to approximate the CDF because otherwise the original binomial CDF is too hard to compute with your TI-84s (because you have one giant factorial divided by another giant factorial and the numerical overflows will kill the computation unless you are very careful about how you go about computing the result).
My favorite go-to guide remains the excellent Calc & Stat Dover book ( http://www.amazon.com/Calculus-Statistics-Dover-Books-Mathem... ), which combines Calc & Stats from page 1. There is simply no better way to learn stats than via calc.
From what I can tell (and remember), elementary & high school math is specifically designed to take you from 0 to calculus. (well, maybe not ALL the way unless you take AP math)
Personally, I find basic stats and prob far more valuable in day to day life than calculus. So my point was just that I wish schools would focus on that area of math as the goal.
But with a mere background of high school algebra, you can learn more about Stats than most college graduates have, and that knowledge is far more relevant to the day-to-day lives of the average person in America than Calculus is.
Those who get a chance to follow through with calculus and apply it do much, much better.
I think the original poster was thinking high school and a target much more like Freedman's text (http://www.amazon.com/Statistics-4th-David-Freedman/dp/03939...).
And what Freedman's book does probably better than any other text in the field is teach how to think about statistics. It doesn't have a lot of formalisms, but if can come to an understanding of what he teaches in that book you'll have a rich understanding of stats.
With that said, if Calc was taught in the context of functions and probability, as in the Gemignani text then I think we'd be better off than how Calc today is focused around engineering.
Of course, anyone beyond the base level of wisdom in this field understands this. It just annoys me that people attempt to diminish the value of statistics with an argument like this.
This is more a warning to people without an understanding of statistics, because most people out there do not have a deep grasp of the fact that correlation does not imply causation.