Why isn't everything normally distributed?
johndcook.com
johndcook.com
There's actually a joke in the field that when you get a new dataset, the first thing you do is fit it to a power law. If that doesn't work, you fit it to a broken power law.
Macro and financial economists have traditionally been the worst offenders.
Mind, you might find different distributions when solving diffusion equations (the Cox-Ingersoll-Ross process involves a Bessel distribution if I'm not mistaken). But the diffusion-jump paradigm is much better justified there than distributional (or even finite moment) assumptions in discrete land.
Put it this way: if finance is abusing distributional assumptions, there's money being left on the table.
Look, even if we disregard jumps as a possible term in equations, merely considering volatility to be stochastic and driven by a Brownian already gets you variances that may grow arbitrarily fast. This is without considering stranger nonlinearities on the Brownian term itself.
People are overly impressed by forceful arguments of the Taleb variety and log-log plots of empirical distributions and start parroting talking points about heavy/fat tails.
If I wasn't computerless and on my phone I would link to a paper that does the analog of the Anscombe quartet for log-log distribution plots "proving that data is Pareto/power law/etc." It's good vaccine for fat tail hipsterism, look it up.
The Tanaka equation is an example (admittedly: not a diffusion, not a case o plug-and-chug the Ito lemma) of a process driven by a Brownian with a discontinuous probability distribution. How the hell? From memory,
dTNK = dB if dB>0 else -dB
Now imagine a model with three equations.
X1 is a bog-standard geometric diffusion (we could have picked something that's chi-squared distributed driven by a Brownian from standard interest rate models) as in the 1970s Black-Scholes models. But instead of having an exogenous volatility, it has its Brownian term dB1 multiplied by a second equation, X2.
dX2 could be a standard mean-reversion equation, but again its dB2 is multiplied by X3.
dX3 is something like abs(dB3 - dB4* X3).
Voilà, an equation (X1) with a sudden break driven by the level of a mean-reverting equation (X2, which tells us volatility should come down in finite time even if it grows by a lot at times) that's set to blow up at a X2-dependent but stochastic level.
Don't get me wrong, Poisson-like jumps are very common (they're precisely the limiting process for sudden jumps) but people overstate (perhaps because they didn't really read the conditions for Ito isometry) how much a Brownian motion forces a system into normality or smoothness.
But hey, people get away with being hipsters about programming languages, why shouldn't they do that for stochastic calculus too, you know?
And unsurprisingly there are quant funds and prop trading firms that use this very fact to make lots of money. Academic economist's love for the normal distribution is derided by pretty much any fellow real world trader I know.
What you're saying is probably part of it, although in my experience that criticism can be leveled as much, if not more, at wet-lab-type biologists who eschew all but the most minimal stats.
With the social sciences, though, there's another phenomenon at play, which is that the phenomena are so abstract often that there's not really a good theoretical reason to assume anything in particular. And if that's the case, because the normal is the entropy-maximizing distribution, you're actually better off assuming that rather than some other distribution. You could also use nonparametric stats, but that has its own advantages and disadvantages.
Bias-variance dilemma and all that.
The truth is, it's hard to beat the normal even when it's wrong. And if you subscribe to the inferential philosophy that every model is wrong, you're better off being conservatively wrong, which implies a normal.
I'm not saying everything should be assumed to be normal. But unless things are (1) obviously super non-normal, or (2) you have some very strongly justified model that produces a non-normal distribution, you're probably best off using a normal if you're going to go parametric. And I think those two conditions are met much more often than we like to admit.
The normal distribution is kind of over-maligned, I think. I started my stats career being enamoured of rigorously nonparametric stats, and still am (esp. exact tests, bootstrapping/permutation-based inference, and empirical likelihood), but have grown to strongly appreciate normal distributions (or whatever maxent distribution is appropriate).
With data in hand, a skew/kurtosis scatter plot is a good way to gauge the higher dimensional distribution of your data. Another option is to cluster the variables of the data set using something like HDBSCAN and color the plot points based on cluster membership.
If you have to go guessing distributions without evidence, you're better off choosing a low k student's T distribution (for robustness to outliers) or a gamma distribution (if you think your data might be skewed).
So there is a huge bias among researchers to assume them to make their treatment of the data easier.
Well, how about the magnitude of stars. So, is the magnitude of stars normally distributed? I googled "distribution magnitude of stars" and got images (by clicking on images tab) that seem pretty normally distributed to me. Aren't they?
That doesn't tell you whether the magnitudes are normally distributed though. That just tells you how the magnitude is computed from an object's observed/measured flux.
http://www.astro.yale.edu/astrom/spmcat/histov.html
Not really. Plus keep in mind magnitude is a logarithmic measure. Things get skewed when you take logarithms. Also if you look for example at star sizes:
http://spiff.rit.edu/classes/phys440/lectures/size/diam_hist...
(from http://spiff.rit.edu/classes/phys440/lectures/size/size.html).
The distribution is certainly not gaussian, and very much likely heavy tail.
Consider the following example: The [Cauchy distribution](https://en.wikipedia.org/wiki/Cauchy_distribution) is a pathological example of a distribution with very heavy tails. Heavier than the power law distributions meantioned in the GP. So heavy, the distribution doesn't even have a mean value. But if you just look at the pdf, it looks almost exactly like a gaussian.
E[(X-EX)^2]
If EX = infinity but X is finite with prob > 0, the standard deviation will be infinite too.
Check out Chebyshev's inequality.
Just don't start talking about it to your pointy-haired boss. You'll get fired.
I'm pretty jaded at this point.
Too bad inference with long tailed distributions is so hard.
The size of a ruler that a machine cuts (quasi normal/Gaussian)
Wealth (Power/Pareto)
Could you elaborate a little, or maybe give an example?
Well worth the 20 minutes.
I recommend reading them through, but slide 39 is the tl;dr.
But if something can happen at any given scale, the distribution must operate very differently. The distribution can't pick out any favorite size to work at --- some fraction of events must happen at every possible scale. So, if we're looking at the distribution of stellar luminosities, then there will be some stars which are as bright as the Sun, some which are 1% as bright as the Sun, and some which are 100 times as bright as the Sun. And since the star formation process doesn't have any preferred scale, then the number of Solar luminosity stars to stars which are 1% as luminous as the Sun must be about the same as the number of stars which are 100 times as bright as the Sun to the number of Solar luminosity stars. it's perfectly fine for the distribution to prefer dimmer or brighter stars because this doesn't end up introducing a preferred scale --- it just tilts the slope of the distribution (which is a line on a log-log graph).
I'm not a mathematician, so I'm not sure if these statements are true in any rigorous sense, but in practice scale-free processes tend to produce power laws in the universe.
It shouldn't, because there is a big reason to lean towards normal distribution, and that is the central limit theorem. If you sample a variety of different distributions, the sum of all these samples will go towards a normal distribution. That is why the normal distribution pops up in so many places.
I've had debates that not everything comes out as a Bell Curve only to be called an idiot for claiming that.
Of course in reality, at some point, some other physical process starts to dominate and cuts off the power law. this then turns it into a power law with a different slope or an exponential cutoff.
Edit: I forgot to add that a star's light profile on an image is actually approximately Gaussian. But it isn't exactly Gaussian --- I seem to remember that the core more closely resembles a Lorentzian. I am obviously not an observer, otherwise I would have thought of that immediately!
[1]: https://en.wikipedia.org/wiki/Maximum_entropy_probability_di...
The maximum entropy distribution with a given mean is the exponential. I think to define power law distributions you will need additional constraints more concrete than "undefined variance".
> it was made after the discovery that on a log log plot everything is a straight line(o)
i'd recommend the entire lecture, but definitely at least check out the anecdote this joke bookends
it is about hubble's original 1929 embarrassing failed attempt to calculate the rate of the expansion of the universe(i)
This sounds wrong. One is invariant under scaling, the other under translation.
It is.
However there is a way to make it correct. Your second sentence is spot on the intuition. If you log transform power-law distributed random variable the RV that you get is an exponentially distributed RV.
That is indeed true, but why should such a property imply that the use of Normal distribution is appropriate ? Just a rhetorical question, of course, because your comment does indicate that Normal is not a good choice unless one has compelling reasons to do so.
Another argument that is used to justify its use is the central limit theorem. That says a sum of (nearly) independent variables with (nearly) identical distributions with finite mean converge to the Normal distribution. If the process under observation is indeed a superposition of such random processes, then yes the choice of the Gaussian can be justified. But it is surprisingly common that one of the 3 requirements are violated. A common violation is that the variance is infinite or it is so high that the process is better modeled as one with an infinite variance. In such situations the family of stable distributions are the more appropriate choice.
Gauss' own use of the Gaussian was motivated by convenience rather than any deep theory. Even at that time it was well known that other distributions, for example, Laplace distribution works better.
The negative log of the (unnormalized) Gaussian density is basically x^2 and this fact comes very handy.
It's used extensively in Bayesian regression, modeling noise as additive Gaussian leads to the squared loss function, and if you also model the prior as Gaussian it leads to an L2 regularizer.
It also simplifies many other calculations. Fitting a mixture of Gaussians by expectation-maximization or the Kalman filter come to mind.
For squared loss almost every theorem you want to be to be true are actually true. In a way its dangerous because it sets up bad expectations (* no nerd pun intended).
A result that holds for squared loss but not for others is the SVD as a low rank approximation of a matrix.
*Expectation minimizes the squared loss
The ONLY variables that are normally distributed are those that are averages of many independent, identically distributed variables of finite variance.
Thus, if you cannot find the finite-variance variables that average up to form a variable X, then X is not normally distributed.
(The parts about independent and identically distributed are technical red-herrings. The only essential condition is finite variance.)
There's something deeply satisfying using a model so simple to explain a piece of nature. :)
Example: cosmic ray energy spectrum. http://iopscience.iop.org/1367-2630/12/7/075009/downloadFigu...
Power law fits to empirical data are heavily misused.
http://cs.unm.edu/~aaron/blog/archives/2007/06/power_laws_an...
"So You Think You Have a Power Law — Well Isn't That Special?"
That's funny -- my field has the same thing with a different distribution. Analyzing stochastic processes is much easier if you assume an exponential distribution. It's one of my criteria for whether somebody's giving a bad job talk. If an unfounded assumption that waiting times are exponentially distributed shows up in the first three slides, the rest of the presentation is probably B.S. Even if the presenter used Beamer and filled the slides with beautiful equations. (The exponential distribution of waiting times is basically an assumption that events are independent.)
Edit: Especially if the presenter's slides are full of beautiful equations.
This is only true if events are independent. Suppose we're modeling rider arrivals at a bus stop. The bus comes once an hour. Does the exponential distribution adequately model how long you need to wait for another person to show up? If a Poisson process is an appropriate model, the expected value of the number of people arriving in the 20 minutes after the bus departs is the same as the expected value of the number of people arriving in the 20 minutes before. Clearly not. Rider arrival rates depend on how long it is until the bus is supposed to depart. Indeed, most arrival processes are not ideal Poisson processes, but it may be a better approximation in some cases than in others.
We use Poisson processes for two reasons:
1. They're often a good enough approximation of reality to make useful predictions.
2. They're mathematically tractable.
Unfortunately, our reasons for using Poisson processes are often more (2) than (1).
None of this before/after bus mess.
Evidently. I'm afraid we have no common ground on which to discuss stochastic processes. Have a pleasant weekend.
Power law walks into a bar. Bartender says, "I've seen a hundred power laws. Nobody orders anything." Power law says, "1000 beers, please".
Agh, I'm a mathematician and should get this, but I don't. Is the joke just that power laws give behaviour that is very small for quite a while, and then becomes suddenly large?
I am a (also former) physicist and it drives me crazy when people fit what ever happens to be in the chart to a straight line, "to get the trend". Whenever I ask them for the theory which predicts that x and y will be linked by y = ax + b, they do not have any.
This also goes on with extrapolations or interpolations, usually without the slightest theoretical reason to do so.
The problem is that when working with marketing, HR or even finance, they are so used to "getting the trend" with a linear fit that I am hopeless.
I once draw a parabole and asked for the trend (between the minimum). They were surprised by this stupid question as "it is obvious that there is no trend". Taking some random points linking share value with the number of women in a company "obviously fits to a trend line".
As Rand Wilcox reports, "Why did Gauss assume that a plot of many observations would be symmetric around some point? Again, the answer does not stem from any empirical argument, but rather a convenient assumption that was in vogue at the time. This assumption can be traced back to the first half of the 18th century and is due to Thomas Simpson. Circa 1755, Thomas Bayes argued that there is no particular reason for assuming symmetry, Simpson recognized and acknowledged the merit of Bayes's argument, but it was unclear how to make any mathematical progress if asymmetry is allowed." (Wilcox, p. 4)
Wilcox, R. (2010). Fundamentals of modern statistical methods: Substantially improving power and accuracy (2nd ed.). New York, New York: Springer.[1]
The reason for this is that if you have the sum independent identically distributed (I.I.D.) random variables (R.V.s), if they converge to a distribution, that distribution is Levy Stable [1], which is power law in it's tails. The Gaussian is a special case in the family of Levy Stable distributions.
The article states that "the sum of many independent, additive effects is approximately normally distributed" which is patently false. The sum of many independent random variables with finite variance is normally distributed. Once you relax the finite variance (and in more extreme cases, finite mean) power laws result.
There are other ways to generate power laws, including having killed exponential processes [2]. There are many other references that talk about the rediscovery of power laws [3] and give many ways to "naturally" create power laws [3] [4] [5].
The article claims that multiplicative processes lead to log normal distributions. I've heard that this is actually false but unfortunately I don't have enough familiarity to see how this is not true. If anyone has more insight into this I would appreciate a link to an article or other explanation.
[1] https://en.wikipedia.org/wiki/Stable_distribution
[2] http://www.angelfire.com/nv/telka/transfer/powerlaw_expl.pdf
[3] https://arxiv.org/pdf/physics/0601192v3.pdf
[4] http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.122...
[5] http://www.angelfire.com/nv/telka/transfer/powerlaw_expl.pdf
Transforming the random variable is a very different thing than transforming the probability density function directly.
The distinction between a random variable and a distribution (or density) is not made clear enough in classes I think.
R = r + e
R^{3} = (r+e)^{3} = r^{3} + 3r^{2}e + O(e^{2})
Unless you expect the error to be significant with respect to r, you can ignore the higher order terms. And voila, you get an approximately normal distribution.The cube of a normal distribution is indeterminate.
Had the relationship between the mass and the radius been linear (M = a * R + b) then yes the mass would have followed a normal distribution as well (with different parameters of course).
That looks a lot like a normal distribution even if it's not one.
$ m = pi * r^3 * p $
m = πr³ρ
(with no disrespect to your use of TeX)
But also it should be m = (4/3)πr³ρ because the volume of the sphere is (4/3)πr³, not πr³.
m = (¾)⁻¹πr³ρ
What happens if they do quality control on both mass and radius?
I imagine the reason why, in your example it is the radius rather than the mass that is normal, is due to what QA is focused on.
If you did QA on mass, you could easily calculate what mass would be required for a particular radius, but that would let misshapen bearings through.
You could do QA on both mass and radius, but unless you have potential contaminants, the radius gives you the mass for free so there isn't much point.
So you might say that the radii = r + e, where r is constant and e is normal around 0 with variance way, way smaller than r. But then the volume of the ball bearings = K * (r+e)^3, and because r >> e, the largest source of randomness term is going to be 3Kr^2e, which makes the whole thing look pretty normal.
Personally, I don’t find it surprising that not everything is normally distributed. Why should any real phenomenon follow a theoretical limiting distribution anyway, never mind a symmetric, infinite-tailed distribution that is exact only in an unachievable limit? The surprise is that so many things _are_ sufficiently near normality for it to be useful!
See
https://en.wikipedia.org/wiki/Central_limit_theorem#CLT_unde...
and
https://en.wikipedia.org/wiki/Central_limit_theorem#Lyapunov...
Even as a quant Taleb was prone to smushing over details to push a narrative. In the 90s already Derman was enabling Taleb into claiming the Black-Scholes formula was an interpolation algorithm already known to option traders and served basically to justify using the risk-free rate as a drift (price trend) parameter in accordance to the "economics establishment". But (as noted by more than one response in the literature) making a different assumption on the drift (or even the distribution) -- i.e. leaving the "Black-Scholes world" and merely interpolating two world-states -- you're left with calibrating a stochastic discount rate that gives put-call parity. But hey, not that technicalities should get in the way of a good story!
Because under mild assumptions, for arrivals, say, of visitors to a Web site, gas station, hospital, the arrivals form a Poisson process, the times between arrivals are independent and identically exponentially distributed, and that's not normally distributed.
Now at nearly any server farm, it is easy to get wide, deep, rapidly flowing oceans of data on the performance of the server farm, and, as U. Grenander once explained to me in his office at Brown, the data is wildly different from what statistics was used to, e.g., in medical data. In that ocean of data, finding anything normally distributed will be very rare.
The claims in the OP about many effects are nothing like good evidence for the central limit theorem or normally distributed. E.g., from the renewal theorem, many examples of Poisson processes are from the results of many independent effects.
E.g., the usual computer based random number generators return what look like independent, identically distributed random variables uniform on [0,1], and that is not normally distributed.
The question in the OP about why not normally distributed is, in one word, just absurd.
https://en.wikipedia.org/wiki/Pareto_distribution#Applicatio...
This makes me wonder if in countries where these environmental effects are mostly optimised (everyone having access to good nutrition), or at least nearly identical for everyone, the normal distribution of height breaks down.
Is height normally distributed in the tallest countries in the world? What about the shortest?
edit: I'll just copy this question to the comments under his blog, maybe the author has some idea about that.
edit2: just noticed the blog post is from 2015... oh well, it was worth a shot.
Check from slide 32.
As someone who doesn't have significant experience in statistics, I'd be grateful for an expert's opinion on the arguments presented in this presentation.
1 - If the problem isn't actually that complicated, the CLT doesn't do much.
2 - If the problem is dominated by one component, it will still mostly look like that component.
3 - Most ways of slicing a normal distribution lead to other distributions. For example the Rician distribution.
Take word distribution in any human language for instance. Word frequencies follow a Zipf distribution because it decreases entropy and hence is more efficient.