Let W be a set of words in a large text. Let f(w) be the frequency of the word w (say, as a ratio between w's count in the text and the most popular word's count in that same text). Assume for simplicity that no two words have precisely the same frequency.
Then we can take all n words of W and put them in descending order by their frequency, thus giving each word a unique index:
w_1, w_2, ..., w_n
where f(w_1) > f(w_2) > ... > f(w_n).
The only information we've organized is an ordering of words by their frequency. We can't really decide in any generalizable way by how much it decreases, at what rate, or anything. (We could, of course, attempt to characterize these things for a single, specific text.) It could be that in your favorite book, the frequency decreases by 0.0000001% for each index in the list for all we know, or it could be that in all Hacker News posts of 2021, f(w_(i+1)) = f(w_i) / 2,
that is, each word in the list occurs half as often as the word just before it. It seems reasonable to believe that if there is some relationship, it should depend on f (which depends on the text being analyzed).What is surprising is that it doesn't really! For all practical purposes, regardless of text, f(w_i) can be written in terms of a simple function of i, specifically as (roughly) 1/i. That means w_1 occurs the most, w_2 occurs half as often as w_1, w_3 occurs a third as often as w_1, w_4 a fourth as often as w_1 (and thus half as often as w_2), etc.
Fairly random. There is no obvious reason to think that words frequency across different languages would be so similar when ranked. Why is it not the case that in language A the 2nd most popular word is used 90% as often as the first, and in language B the 2nd most popular word is used 70% as often as the first?
One would think with something like languages which are so complex and developed independently that there wouldn't be such a surprisingly consistent ratio.
And beyond language, it occurs in stuff city populations, the amount of traffic websites receive, last names, ingredients used in cook books, etc...
1. 4096
2. 2048
3. 1024
4. 512
5. 256
6. 128
...
Here, the frequency of a word is inversely proportional to its rank to the power of two, not the rank itself.
An interesting application is looking at untranslated works -- for example while they can't translate the Voynich Manuscript, it does follow the distribution, so it is probably not just random scribbling (of course, this doesn't rule out the likely options of cypher or constructed language).
Of course the frequency is going to be proportional in some way to the rank. But there are many ways that could happen. #1 could occur 10% more than #2. Or twice as much.
And for the law to hold true no matter how deep you go is also surprising. Language seems like it should be a little more chaotic than that, with the top, say, 50 words following one distribution, then the longer tail kinda bumping around at different slopes.
This is my lay understanding. Corrections welcome :)
For comparison, Zipf has a power law tail (x^{-\alpha}) whereas a Gaussian distribution falls off doubly exponentially (e^{-x^2}).
Shows up "everywhere", not exclusively. For that matter, also Erlang distributions are occurring models.
That's because it's quite possible to have 2 lines of fit that look like they both very nicely follow a power law, but then once you go from a log-scale plot to linear plot by doing the inverse of the log, ie exponential transform, one of the 2 plots may turn out to be really crappy, actually.
So, if the error bars on the log-scale plot are not in some sense really small (e.g logarithmic in themselves) already to begin with, then, you may actually be committing ~great crimes of statistics~, unknowingly.