For example, you might have a radio signal (such as WiFi) that you want to receive. First step is that you have to pick that signal out of whatever signal comes out of your radio receiver: which will be the WiFi signal along with all sorts of noise and interference from other users. Typically the search will be done with the mentioned "Pearson's Correlation", using it to compare the received signal with an expected template: a value of 1.0 meaning the received signal is a perfect match with the template, a value of 0.0 meaning no match at all. If the wanted signal is present, interference, noise and distortion will reduce the result of the correlation to less than 1.0, meaning you might miss the WiFi signal, even though it is present.
This article is about coming up with a measure that gives a more robust result in the face of noise, interference and distortion. It's fundamental stuff, in that it has quite general application.
Skimming it now, this looks wild. Using the variance of the rank of the dataset (for a given point, how many are less than that point) seems... weird, and throwing out some information. The author seems legit tho, so I can't wait to try drop-in implementing this in a few things.
The neat thing about ranks is that, in aggregate, they're very robust. You can make an estimate of the mean arbitrarily bad by tweaking a single data point: just send it towards +/- infinity and the mean will follow. The median, on the other hand, is barely affected by that sort of shenanigans.
1. a mutual relationship or connection between two or more things
2. [Statistics] interdependence of variable quantities.
3. [Statistics] a quantity measuring the extent of the interdependence of variable quantities.
The most sympathetic to your definition is Wikipedia: In statistics, correlation or dependence is any statistical
relationship, whether causal or not, between two random variables or bivariate
data. In the broadest sense correlation is any statistical association, though
it actually refers to the degree to which a pair of variables are linearly
related.
And that's the mathematical formulation. Correlation also has a meaning in everyday speech, and mathematics doesn't have the authority to just adopt terms and then claim people are wrong after they've changed the meaning.Also correlation very definitely means that knowing <x> tells you something about <y>. And vice versa. Like, for example: its value. Or at least a better idea of it than pure guessing without correlation.
Correlation, in general, just means some sort of statistical dependence: knowing x tells you something about y. It's often "operationalized" by computing Pearson's r: it's easy to do and there's lots of associated theory.
However, I would find it absolutely bizarre if someone showed a plot with obvious non-linear dependence and described it as "uncorrelated". In that case, the low r reflects a failure of the measuring tool rather than something being measured.
Worth noting the author is a highly regarded professor at Stanford.
I'd like to learn more about the small sample properties. Proofs of asymptotics are necessary but less interesting. But the author's examples on example data sets look like it makes sense.