Avoiding a Common Mistake with Time Series
svds.com
svds.com
It looks like what the author is really trying to say is that we should pass the data through a high-pass filter, eliminating any 'expected' trends such as inflation, and instead observing if the noise of the two datasets is correlated. This is an observation that has some value, but is certainly not trivial to pick a threshold for the high-pass (certainly it's not always just a linear trend), and the mutually dependent variable can have as much noise (if not more) as the two measured data sources, so you still might get "false" correlation.
Causation is super tricky business. You probably have an okay model of it intuitively once you've done a little stats work, but to explain it to someone else or tackle it in more exotic or hypothetical domains is an exercise is insanity if you're not armed with the very best tools.
Of course what you say is true, but executing it in practice and having a sufficiently rich language to identify confounders well is what's usually lacking.
His statement (after adding the constant trend) is misleading:
"Now let’s repeat the same tests on these new series. We get surprising results: the correlation coefficient is 0.96 — a very strong unmistakable correlation"
What he's calculated is the correlation between a set of points from y_1 and y_2 - and that will be large, of course - their (deterministically increasing) means have correlation 1.
The quantity that qualifies as correlation for predictability purposes is actually the correlation between the deviations from mean.
This is all fairly clear if you use the actual formula (Pearson correlation):
E[(y_1 - mu * t)(y_2 - mu * t)] / (sigma_y1 * sigma_y2)
I believe the beauty is you're using time (x and x-1) to preserve the change from sample to sample but remove the time influence that broadly affects all and distorts the subtle differences that may in fact not be related.
It only matters if the commonality is far more influential and unimportant than the unique but import parts.
So, in effect, you model the correlation between CO2/time vs. temperature/time. The interesting correlation would be if every time CO2/time deviates you can find a correlation with a temperature/time deviation.
Put another way, we've introduced a mutual dependency. By introducing a trend, we've made Y1 dependent on X, and Y2 dependent on X as well. In a time series, X is time. Correlating Y1 and Y2 will uncover their mutual dependence — but the correlation is really just the fact that they're both dependent on X. In many cases, as with Jennifer Lawrence’s popularity and the stock market index, what you’re really seeing is that they both increased over time in the period you’re looking at.
What, pray tell, is the X for which the stock market index and Jennifer Lawrence's popularity depend on? Oh, you say they're both dependent on the underlying trend... What?
Edit: I'll note that this is the same thing as subtracting a very long-windowed low-pass filter, i.e., performing high-pass filtering.
Doing a very quick A/B test helps too.
In my opinion, the absolute, utter core essence of science can be expressed simply as "Always be trying to prove yourself wrong." The human brain is extremely biased in the other direction, and it's darned good at proving itself correct. It can prove itself correct in absurdly powerful ways. Always be fighting it, always be looking for ways to prove yourself wrong. If you do it seriously, things like "the scientific method" will naturally fall out of your serious attempts and need not be surrounded by near-worship, whereas no amount of worshipfully-following a checklist of the "scientific method" without the true effort to prove yourself wrong will produce truth; the human brain is far more powerful than the "scientific method" or any such static methodology.
I have deliberately left the word 'theory' out of this post. This scientific mindset beyond that into engineering and any number of day-to-day activities.
Insanity!
I agree with you for the most part. But there's a line. When it starts to destroy your personal identity and sense of self, you've crossed that line. Sometimes it's just better to be right and trust yourself.
Or not. That's zen, I think. Trying to have so much information about everything that you wind up overloading yourself with it and can't make heads or tails of it, because in every simple truth exists an infinite proof. Analysis paralysis!
Alas, if you're looking for an excuse to tie either yourself or me up in some sophomoric philosophical conundrum, this doesn't do it. But there's no lack of such things if you look, so don't be disappointed. Keep on trucking. May I suggest Godel, Escher, Bach: An Eternal Golden Braid if you would like the industrial strength version rope, err, braid to tie yourself in?
But I'm sure there are two ways to respond to my comment: pretend you know what you are talking about, or admit you could be wrong and care to engage me in a real conversation. I've been down the same road hundreds of times on the internet, and it is hard to find people to talk to about things and meet in some kind of scientific, logical middle. But if that's not your thing, oesn't bother me if all you came here for was to prove yourself right about the noble ethics of your science. At least I'm no longer trapping my mind in a paradox of it's own creation. (Or am I? I never really know).
Logic is very different from science, that's all I have to say. Rigor is for precision, not discovery.
Considering the entire rest of your message consists of you essentially trying to lord over me whatever superiority you seem to think you have, and not missing an opportunity to slide a snide comment in, which I'd also observe is consistent with your original post, I've come to the conclusion that the rest of your post belies this claim. You're being abrasive and abusive while trying to pretend to be so much more open minded.
No sale. I recognize that tactic and refuse to engage with it. I'm sticking with my original assessment; you're being too sophomoric to engage with, and I reject your attempts to psychology me into engaging with you in the presupposed frame of your superiority.
I don't like psychology.
But then there are all the ways people screw up A/B tests...
A famous example of this:
The tale of David Leinweber, which is related in the excellent new book "Quantitative Value," illustrates this point about "stupid data miner tricks." Leinweber sifted through a United Nations CD covering the economic data of 140 countries. He found that butter production in Bangladesh explained 75 percent of the variation of the S&P 500 Index. Not satisfied, he found that if he added a broader category of global dairy products, the correlation would rise to 95 percent. Then he added a third variable, the population of sheep, and found that he had now explained 99 percent of the variation in the S&P 500 for the period 1983-'99.
(http://www.cbsnews.com/news/what-butter-production-means-for...)
I reckon they had a look at the competition before putting it in the envelope
So a small correlation (e.g. r=0.10) can still be "statistically significant" at p<0.001 but all this means is that r is reliably different than 0.00 --- it doesn't mean r is big