Zipf’s Law Arises Naturally When There Are Underlying, Unobserved Variables
journals.plos.org
journals.plos.org
This is not discussed in the associated paper either, though the paper concludes that the distribution is a good representation of common Chinese words (i.e., it isn't skewed). Any clues to why that is?
[0] http://crr.ugent.be/programs-data/subtitle-frequencies/subtl...
I don't know why the word distribution doesn't follow a power law, but I can tell you that it isn't a good representation of common Chinese words. I downloaded the data set a while back, intending to use it for improving my vocabulary. I found some words (like the Chinese word for 'vampire') had too high a frequency.
The lists are based on American TV shows' Chinese fansubs. I guess vampires and zombies are common themes.
I don't consider myself an expert in word-frequencies, but the article makes a convincing argument that these frequencies are representative of natural Chinese and a better or just as good measure as their baseline corpus.
The joint probability distribution of states of many degrees of freedom in biological systems, such as firing patterns in neural networks or antibody sequence compositions, often follows Zipf’s law, where a power law is observed on a rank-frequency plot. This behavior has been shown to imply that these systems reside near a unique critical point where the extensive parts of the entropy and energy are exactly equal. Here, we show analytically, and via numerical simulations, that Zipf-like probability distributions arise naturally if there is a fluctuating unobserved variable (or variables) that affects the system, such as a common input stimulus that causes individual neurons to fire at time-varying rates. In statistics and machine learning, these are called latent-variable or mixture models. We show that Zipf’s law arises generically for large systems, without fine-tuning parameters to a point. Our work gives insight into the ubiquity of Zipf’s law in a wide range of systems.
My quick understanding of this paper is that lurking variables can cause Zipf's law. If enough variables are controlled for, in some domains Zipf's Law goes away? For example, they included parts of speech for natural language and for some parts of speech Zipf's Law wasn't existent but for others it was.
This makes intuitive sense to me.
They also find a way to estimate how much of a Zipf's Law effect is caused by a particular variable. Seems like a nice test to run in Stata or R.
Most probability distributions are designed to model some kind of process. The time between clicks of a geiger counter is exponential, the number of clicks in a second is poisson. The amount of time your geiger counter lasts before it breaks and you have to buy a new one is weibull. That's the whole reason why we invented probability distributions in the first place: to model processes.
The class of distributions that both power law tails and the Normal distribution fall under are called "Levy-stable distributions" [1]. Given the sum of independent, identically distributed (I.I.D.) random variables that converge to a stable distribution, the limiting distribution is Levy-stable (read, power law tails). If you the sum of the I.I.D. random variables has finite variance, the Gaussian distribution is the limiting distribution, which is a special case of the Levy-stable distributions. If each of the I.I.D. random variables has infinite second moment (e.g. infinite variance) and converges, then it converges to something with power law tails.
We've been spoiled to think that the Gaussian's are the only or most common limiting distribution (thus the name "Normal") but Levy-stable distributions are another class of "common" distributions for the same reason: a common limiting distribution of sums of (independent) random variables have a limiting distribution that has power law tails. This was one of Mandelbrot's work horses in that power laws show up in nature, are self similar and represent a kind of "statistical" fractal [2].
There are other ways to generate power law tailed distributions as well, including having a random variable that has an exponentially increasing value with an exponentially decreasing probability of occurring, where each go at a different rate, say [3].
[1] https://en.wikipedia.org/wiki/Stable_distribution
[2] http://www.webofstories.com/play/benoit.mandelbrot/32;jsessi...
Is this just a complex way of saying "when you put items in ascending order, that order's the inverse of the same items listed in descending order", or have I missed something?
Thus the most frequent word will occur approximately twice as often as the second most frequent word, three times as often as the third most frequent word, etc.: the rank-frequency distribution is an inverse relation.
equivalently
a plot of rank v frequency is a hyperbola, which implies a log-log plot is linear.
Disclaimer: I'm not an expert, I looked at the Wikipedia page in response to this story. https://en.wikipedia.org/wiki/Zipf's_law
rank * frequency = constant