Modes, Medians and Means: A Unifying Perspective
johnmyleswhite.com
johnmyleswhite.com
The generic question is: Given a loss function, what "property" of the distribution minimizes average loss; and given a "property", characterize all such loss functions. For example, Bregman divergences are (essentially) all losses that "elicit" the mean of a distribution. If you have any monotone continuous function g(), then |g(x) - g(s)| actually also elicits the median, and these are essentially all that do.
Apologies for self-promotion, but you can read more at references on this page (disclaimer: I'm one of the researchers who posted it): https://sites.google.com/site/informationelicitation/
or tutorials on this subject at my blog: http://bowaggoner.com/blog/series.html#convexity-elicitation
* L_0 -> mode
* L_1 -> median
* L_2 -> mean
* L_infinity -> midrange, I think
that is, (smallest observation + largest observation)/2
(BTW, the author is also a big contributor to the wonderful Julia language, I believe)
Wikipedia entry: https://en.wikipedia.org/wiki/Norm_(mathematics)#Maximum_nor...
For what it's worth, I think the confusion of your comment and jules's below comes from the idea that the lp-norm is sum(x_i^p), when it's actually [sum(x_i^p)^1/p] (please forgive my notation). Since raising to the 1/p power is monotonic, it doesn't actually make a difference in minimizing or maximizing the norm, so people often use them interchangeably.
[0]: http://www.johnmyleswhite.com/notebook/2013/03/22/using-norm...
We assume that for one dimension, we can tell the length: length of d is just |d|, the absolute value of d. But what if we have many dimensions i=1..N?
Then we can use this L_p norm, and the idea is basically just that we set the norm to
[ sum_i=1..N |d_i|^p ] ^(1/p)
Now, what does that give us? Note a few things:* when all of the d_i are zero except for one, we get back just the value that's not zero (in absolute value), because we go to the p-th power and then back. So, we recover the intuition (and number) from the one-dimensional case.
* for p=1, we just add all the absolute values of the d_i. That just gives us the citiblock/Manhattan norm.
L1 norm of [3, 4] = 7.
L1 norm of [100, 1, 1] = 102.
* for p=2, we square all values, add, and take the square root. This, by repeated application of Phytagoras, gives us the usual "Euclidean" distance in space: if you're 3 m west and 4 m north, then you're 5 = sqrt(9+16) m away.
L2 norm of [3, 4] = 5. L2 norm of [100, 1, 1] = 100.00999
* now, think about what happens if d_1, say, is very big, and all the others are tiny. If we add them (L_1), we just get the sum, a bit bigger than d_1. If we add the squares, and then take the square root, we'll be even closer to d_1, because the square of d_1 is going to be very big, and the other squares really small, and then we add and take the square root to basically get back d_1. See example.
* Now imagine p=4, or 100. We take a huge power of all d_i, then add, then "invert" the power. All the terms in the sum are going to be dwarfed, except the biggest one!
L4 norm of [3, 4] = 4.284572295
L4 norm of [100, 1, 1] = 100.0000005
* Thus, the infinity norm goes towards the maximum of the d_i. And that's what it's defined as: max_i=1..N |d_i|
* So, big p tend to exaggerate the differences between the d_i, and only the biggest counts. Conversely, small p tend to make the d_i more similar, until it only really matters whether a d_i is 0 or not. In other words, you just count how many are not zero. And that motivates the L0 norm.
Now, think about the unit circle in 2D.
Here some ASCII art:
L0:
|
-+-
|
L1: /\
/ \
\ /
\/
L2: -
/ \
( )
\ _ /
Ok, that's my best ASCII rendering of a circle...L_infty:
___
| |
|___|
You see that the 4 points up/down and left/right are fixed (-1, 0), (0, 1), (1, 0), (0, -1), because of this property that if all are zero except one we fall back to the normal distance. But then, for others, the points we consider close enough to be within the unit circle "bulge out" as we increase p.It doesn't even mention variance (second central moment), let alone skew or kurtosis?
https://stat.ethz.ch/pipermail/r-help/2005-December/083875.h...
"pkg:moments uses the ratio of 4th sample moment to square of second sample moment, while pkg:fBasics uses the variance instead of the second moment and subtracts 3 (for reasons to do with the Normal distribution)."
"The "correct" number for kurtosis depends on your purpose. The number for "kurtosis" that subtracts 3 estimates a "cumulant", which is the standard fourth moment correction weight in an Edgeworth expansion approximation to a distribution. Neither of the numbers described ... compute the "4th sample k statistic", which is "the unique unbiased estimator" for that number (http://mathworld.wolfram.com/k-Statistic.html)."
notice that the unifying perspective depends on a continuous parameter p, which is 2 for the mean, 1 for the median and -infinty for the mode. Thus, there is a continuous family of statistics that interpolate between these three things!
he does not mention it neither, but this means that you can define modes of a continuous variable (without resorting to histogram bins)
> There is no unanimous opinion on leaving the value generally undefined.
That is a pretty roundabout way of saying that people agree that it should be _something_, and generally disagree on what that something should be.
Pragmatically, what this means is that 0^0 is not sufficiently specified (just meaningless symbols) unless you prescribe the context/meaning with which you're using it. And with regards to defining summary statistics, the article talks about a context where the exponent scans over different (continuous) values, passing by {2,1,0} along the way.
It should be an interesting exercise to consider what happens when the exponent approaches positive or negative infinity i.e. large magnitude positive or negative numbers.
Mathematically, you don't take limits as a sequence of single-variable limits. You take them by constricting an n-dimensional circle (two points / circle / surface of a sphere / etc.) around the point whose limit you're interested in.
This immediately implies that there are infinite possible approaches to the (0,0) point, rather than only two as there would be if you were taking it as two single-variable limits in sequence.
julia> vecnorm([0.0, 1.0, eps()], 0)
2.0
julia> vecnorm([0.0, 1.0, eps()], 1)
1.0000000000000002Arithmetic difference satisfies for first property
Squaring each difference satisfies the second property
Taking the arithmetic mean satisfies the third property
Var(X) = E[(X - mean)^2]
To get variance, you need to be chiefly concerned with distributions that have a variance in the first place - and then the additional information contained in that statistic has an amount of descriptive power over that of the median.
EDIT: just saw this was mentioned in one of the comments...
https://en.wikipedia.org/wiki/Fr%C3%A9chet_mean
https://en.wikipedia.org/wiki/Mode_(statistics)#Use
https://en.wikipedia.org/wiki/Central_tendency#Measures
https://en.wikipedia.org/wiki/Nonparametric_skew#Relationshi...
https://en.wikipedia.org/wiki/Geometric_median
https://en.wikipedia.org/wiki/Weber_problem#Definition_and_h...
https://en.wikipedia.org/wiki/Centerpoint_(geometry)
https://en.wikipedia.org/wiki/Trimean
https://en.wikipedia.org/wiki/K-medians_clustering
https://en.wikipedia.org/wiki/Medoid
https://en.wikipedia.org/wiki/Generalized_mean
https://en.wikipedia.org/wiki/Quasi-arithmetic_mean
https://en.wikipedia.org/wiki/Lehmer_mean#Special_cases
https://en.wikipedia.org/wiki/Logarithmic_mean#Generalizatio...
https://en.wikipedia.org/wiki/Stolarsky_mean#Special_cases
https://en.wikipedia.org/wiki/Facility_location
See Also:
https://en.wikipedia.org/wiki/Problem_of_points
https://en.wikipedia.org/wiki/Chebyshev%27s_inequality
It unifies the inequalities:
max > root mean square > arithmetic mean > geometric mean > harmonic mean > min
that I remember from highschool math competitions.[2]
[1] http://en.wikipedia.org/wiki/Power_mean#Special_cases
[2] https://artofproblemsolving.com/wiki/index.php?title=Root-Me...
1. The gaussian distribution is important (all sums converge to it) and the squared difference from the mean characterizes it.
2. Euclidean space is important, and squared errors stay the same if you rotate everything. (Other errors don't).
3. Linear regression with squared error has a closed form solution. Other types of models converge very fast when using it because the further you are from optimal, the bigger your gradient is.
With ^1 you get the Laplace distribution. This is widely used when the data is sparse (has many elements exactly 0). This is less sensitive to ourliers and in image reconstruction it gives sharper images than ^2.
If the distribution of losses is skewed or has outliers then estimates other than the mean (median, trimmed means etc) often under-estimate total losses.
Under-estimating total losses in the long run could be very bad for business.
https://qz.com/260269/painfully-american-families-are-learni...
and the numbers being summarized. They only differ in the type of discrepancy being considered:
The mode minimizes the number of times that one of the numbers in our summarized list is not equal to the summary that we use.
The median minimizes the average distance between each number and our summary.
The mean minimizes the average squared distance between each number and our summary."And "distance" is "the 1st power of distance".
And "squared" is "2nd power"
I struggled with this choice myself and so far, decided to do the same as the author with a noscript tag to explain to the reader why I'm loading cloudfare code.
https://gist.github.com/anonymous/2e7972e9282602118be24cd1fd...
In fact, there's a whole spectrum of averages defined with mean and median on each end, depending on how many outliers you eliminate. For example, if you have eight numbers, you can define a spectrum of four averages:
2,3,5,7,11,13,17,19 // mean, here 9.6250
3,5,7,11,13,17 // mean with outlier on each side stripped, here 9.3333
5,7,11,13 // mean of central two quartiles, here 9.0000
7,11 // median (i.e. mean of center two numbers), here 9.0000
You could then repeat the process on that spectrum of averages to get a shorter spectrum, here [9.2396 (mean), 9.1667 (median)], recursively until you have one "mean-median" left, here 9.2031.I wonder how this fits in with the explanation in the post.
I think "removing outliers" is also snugly in the camp of practical heuristics. A mathematical definition might not want to automatically eliminate outliers when "outlier" is also subject to a choice of definition.
This is actually not true!
The correct way to do this would be to take the (right-hand) limit of argminₛ ∑ᵢ |xᵢ - s|ᵈ as d → 1⁺. You get back a unique number and it is not the mean of the middle two values.