My Favorite Statistical Measure: Hoeffding's D
github.com
github.com
This is incorrect.
Hoeffding's D was not intended to be used as a descriptive statistic to measure the strength or direction of a relationship, it's just a statistic from a nonparametric test for independence. It comes from Hoeffding's 1948 paper A Non-Parametric Test of Independence [1]. Interpreting it as a measure of the strength of a relationship is questionable. The scale is... mostly meaningless except in the qualitative sense that close to 0 is close to independence (maybe--more on that later) and close to 1 is a strong relationship of some kind. But what does a D of 0.2 or 0.6 or 0.9 mean? Who knows! It's certainly not your traditional correlation scale--don't be tempted to interpret it that way.
Interpreting it as a measure of the direction of a relationship is simply wrong. You can easily check that D for two vectors X and Y is the same as D for X and (-Y). It gives you no information about the direction of the relationship. D sometimes being negative is just an artifact of the way it's calculated and scaled.
Back to that "maybe close to independence" thing--Hoeffding's D is blind to certain deviations from independence. You can have a joint distribution with clear dependence but Hoeffding's D will be 0. See Section 4 of [2] for some examples.
If you want to use an esoteric measure for the strength of the relationship between two variables I'd go with the distance correlation instead. At least that has a clear meaning.
[1] https://projecteuclid.org/journals/annals-of-mathematical-st...
If you’re using it more as a form of search (information retrieval), then I think there’s no harm in using it. For example, for ranking relevant embedding vectors.
Apparently it works quite well for finding similar genes (I guess you replace base pairs with integers or something like that). Sometimes you just need a good place to look and then you can confirm things independently.
> The final formula for Hoeffding's D combines D_1, D_2, and D_3, along with normalization factors, to produce a statistic that ranges from -0.5 to 1. This range allows for interpretation of the degree of association between the sequences, with values near 0 indicating no association, values closer to 1 indicating a strong positive association, and values near -0.5 indicating a strong negative association.
> And a score near -0.5 suggests they're moving in opposite directions, perhaps clashing rather than complementing each other.
If there's a decent relationship there, you could run Pearson, Hoeffding, Chatterjee, a simple linear regression: it'd be a weird dataset where you get different results.
> Suppose you have two sequences of numbers that you want to compare so you can measure to what extent they are related or dependent on each other.
Great, got it. I'm on board. I have two sequences of numbers.
> [20 more sentences about why you might have two sequences of numbers and what they might look like]
I assure you that anyone interested in an article titled "My Favorite Statistical Measure" doesn't need anything besides that first sentence.
It's totally possible to target both the amateur and expert audience in a single article, but it need an appropriate structure to achieve this :)
For exemple:
(1) Abstract: set the direction of your article. What we want, what we use (Pearson), what I like to use (Hoeffding's D), eventually some outcome
(2) Introduction: your first three paragraph can go in there, with two subsection (exemple, and state of the art with Pearson use).
(3) Pearson details
(4) Hoeffding(s D details
(5) implementation of Hoeffding's D
(6) Conclusion/comparaison
Personally, without a clear introduction stating where we start, where we go, and what is the journey programm, I tend to not read (which is not good for me and for you :) )
https://en.wikipedia.org/wiki/Phi_coefficient
I'm asking this because association measures for binary variables are often easier to analyze.
Although, I guess you could theoretically try taking the values in each sequence in groups of N at a time and then interpreting each grouping of binary digits as an integer, and then compute the Hoeffding’s D of the resulting integers. Maybe doing that for a range of N values, like 3 to 8, and averaging the Hoeffding’s D values you get. Not sure if that really makes sense though, I need to try it with some real data!
Generically, correlation has nothing to do with whether data are related, dependent or independent.
It only is an indication of this *if you already believe* they are dependent. There has to be a prior semantic model of what the data means (what it is a measure of, how reliable its a measure of it, etc.) before correlation measures anything at all.
The only reason we would suppose correlated data to be related, is just that we design experiments to already contain possible dependencies. We do not, as a habit, measure say, the fall of rain and the beat of a song playing. But these would, given suitable measurements, be correlated.
This becomes absolutely vital to understand when experiments do take this "hoover up everything, arbitrarily" approach. As often they do in the social and psychological sciences.