Why 30.1% of numbers start with 1
dspguide.com
dspguide.com
http://falkenblog.blogspot.com/2008/12/benfords-law-catches-...
Predict and then create a histogram of the leading digits of the file sizes of the non-zero-length files on your computer.
[SPOILER: when I did this once I found sharp peaks around digits that weren't 1. You are likely to see this if you have a large number of files around a particular size, examples: a) whatever your digital camera typically produces; b) whatever size your software encodes a typical song into. These files violate the assumption that you are sampling sizes over a wide range of sizes. After excluding these files I observed Benford's law quite closely on the remainder.]
1 ****************************
2 ***************
3 *********
4 *****************
5 *******
6 ******
7 *****
8 ****
9 ***
That spike for 4 is due to the default directory size of 4096 (my experiment included directories as well as files). The information was pulled from 503,444 files and directories.1 29.12 %
2 20.60 %
3 13.99 %
4 12.54 %
5 5.74 %
6 4.78 %
7 4.84 %
8 4.93 %
9 3.47 %
Your files must have been made up! ...or you have a nice demonstration of how people shouldn't do too careful calculations with Benford's law. "30.1%" — 3 significant figures, really?
cd /
find . -exec stat -f "%z" {} \; | cut -c -1 > /tmp/tally.txt
sort /tmp/tally.txt | uniq -c
Mine came out with...
506 0
80370 1
30396 2
25215 3
21787 4
22174 5
26251 6
12810 7
10455 8
5556 9
Very interesting...find . -type f -exec stat -c %s {} \; | cut -c 1 | sort | uniq -c
Note that I exclude directories to avoid the size 4096 bias.
I ran it in my "project" directory and found that 38% of my file sizes begin with "1". That directory includes Perl source code files, input data files, and automatically generated output files.
After the digit "1" the distributions ranged from 3% to 9% with no obvious bias I could see.
find . -type f -ls | awk '{print $7}' | cut -c -1 | sort | uniq -c
Now I'm piping that into Perl to convert the counts to percentages. If I figure out a one-liner for that I'll let you know.
Next I'll be tempted to write a module for generating "realistic" (Benford-compliant) random numbers using this concise specification from HN contributor "shrughes":
"Data whose logarithm is uniformly distributed does [follow Benford's Law]."
I could use that to produce demo or test data.
2192 0 - 0%
389003 1 - 38%
151943 2 - 15%
116663 3 - 11%
96393 4 - 9%
76590 5 - 8%
53572 6 - 5%
45381 7 - 4%
47138 8 - 5%
36983 9 - 4%
http://www.billthelizard.com/2009/04/benfords-law.html
</pedantry>
The author also claims that looking at the Fourier transform of the probability distribution is key to understanding what's going on. But the full extent of his Fourier-based analysis is this: Consider the probability distribution function for log_10(data). Then Benford's law holds if this is constant (editorial note: it cannot in fact be constant) and holds roughly if it's roughly constant. That happens, kinda, when the probability distribution is very broad (editorial note: no, not really; see the example above). What, you didn't see anything about Fourier transforms there? Well, that's because the Fourier stuff is really almost all window-dressing.
For a brief account of Benford's law and related matters written by someone with a better grasp of what's going on, you could turn to http://terrytao.wordpress.com/2009/07/03/benfords-law-zipfs-... whose author is one of the best mathematicians currently living and also a very good expositor.
And the British MPs' expenses: http://www.jgc.org/blog/2009/06/its-probably-worth-testing-m...
And BBC executives' expenses: http://www.jgc.org/blog/2009/06/running-numbers-on-bbc-execu...
I wonder why this law is used for detecting fraud...
In particular: dimensioned measurements must represent the same relationships between data under many different scales thus leading to logarithmic sampling.
Without further information to give expectations to other trends, these laws great starts.
Still, sociologically, working scientists tend to believe in these models as if they were hard rules, to the extent of constructing bridges, rockets, nuclear reactors, nationwide health recommendations and global financial systems without fundamentally understanding why each distribution might arise, and why it might fail to explain real phenomena.
- Wolfram Mathworld references this topic and claims Benford's Law was put on a rigorous footing in 1998 (http://mathworld.wolfram.com/BenfordsLaw.html), but this is not even mentioned in the book.
- the author makes "straw-man" type claims that (unnamed) prominent mathematicians view Benford's law as "paranormal". Also see the last two paragraphs on the first page of the original article, where the author dismisses the idea of a "universal distribution", which is used at Mathworld to give a heuristic derivation (suppose some rule governs this distribution -> apply scale invariance -> derive needed properties) - it seems like he misunderstood this.
- the author claims that he is the first to have solved this mystery, but doesn't reference any literature since 1976
- he claims on a blog (http://www.dsprelated.com/showarticle/55.php) that he tried to publish in journals, but was rejected because mathematicians weren't interested. He then published in a textbook, not even in some sort of paper. His "proof" is really long, uses very elementary mathematics and unnecessary computer programs, and refers back to other parts of his book (so that I don't want to actually try to parse the whole thing and see if I believe it).
Can anyone confirm or deny my suspicions?
That said, Benford's Law is really cool.
Ted Hill's papers on Bedford's law are from 1995, and this chapter is from 1997. Wolfram's date is incorrect, although the secondary phenomona that choosing from a distribution which itself was chosen at random is log normal was proved at that later time.
The 'straw-man' was misunderstood by you. The author points out in the conclusion of that paragraph that all of these pseudo-scientific or grandiose explanations where nonsense.
The 'proof' is really an explanation that the phenomena presents itself to an easier analysis when viewed in terms of FTs. The computer program is there to show the reader the 'repeatedly divide by ten' action which is implicitly going on when we map from the unbounded domain to a small bounded domain.
Its a nice explanation really, but as the problem was solved years earlier and this analysis isn't given with another application, I can see why it wasn't accepted for publication in a pure math journal.
I think 1 is special, in any base, because it's the first digit used when a new digit gets added. If you're talking about quantities that vary easily by, say, thousands, once it crosses the 10,000 threshold the first digit changes much more slowly.
There's more to think about here for sure, though.
[Not regarding you comment or this response]
As I have understood, the Law applies only to a logarithmic scale [1 to 2, 2 to 3, 3 to 4, ... to ...]. Look at this pattern graph:
Try it on a set of numbers in base 10, then convert them to base 16 and check the percentages you get. They still follow the same pattern, but of course there are more slots and the individual percentages are lower because of that.
Eh? Shouldn't this be a uniform distribution?
I wonder whether number choice will fallo Benford's Law.
Now, pair them off, in order, and write down the products: you'll notice a large amount start with 1!
(of course you'll actively avoid this if you're expecting it)
It's pleasing to note that the law applies to itself; Stigler was not its discoverer, and of course he was well aware of this when naming it.
If you saw logarithmic scale ( http://www.ieer.org/log.gif ), you know that distance between 1 and 2 on logarithmic scale is much wider than between 2 and 3, and so on. So, no surprise, if you seed linear data through logarithmic (non-linear) filter, then numbers will follow pattern of logarithmic scale.