[EDIT: I had some other criticisms that were incorrect and overly harsh anyway. I've removed them.]
[EDIT: I had some other criticisms that were incorrect and overly harsh anyway. I've removed them.]
Also, to defend myself a little: I think I'm at least being responsible by in no way claiming that it's statistically valid, and in fact making the point that it's not, and citing the reasons why, several times in the article.
I do stand by my point, however, that this helps the public understand the danger of metadata for data-mining, as well as introducing them to the pitfalls of statistics.
The most important thing is getting an estimate of how good your averages are. Try to modify a couple of things, for example how do the 67% change with sample size (verification)? How does the number turn out if you feed it biased data, e.g. only male phone numbers (falsification)?
The goal is to use this as a hook to get the public interested, then take them for the ride as I learn. That's why I tried to be really cautious about pointing out all the reasons why this particular model is bad.