A Mathematical Analysis Of The Iranian Elections
washingtonpost.com
washingtonpost.com
But let's talk about the real issue. When you forbid independent monitors, announce results before the votes gets counted, and conclude absolute certainty in the results before there is a statistical analysis, it is hard to believe in a fair election.
If you try a whole bunch of fairly arbitrary metrics, e.g. how many end in 7, how many end in 5, how many have consecutive increasing digits, and you mine the data into finding two arbitrary ones that random chance would say occur 5% of the time, does that really tell you anything about the data? I find this article a stunning example of the dangers of mathematical ignorance.
You are of course right, why choose exactly those two metrics? Also, the two metrics are probably not independent, so taking their product is bogus.
The article was written by a 'PhD candidate' in political sciences. I think he should get his degree. He learned how to convincingly misuse statistics.
I'm not familiar with the digit-distribution law the article mentioned, but on its face it doesn't seem entirely unreasonable. I'd prefer to see some study applying this same analysis to thousands of other elections, of course.
But, prima facie, using digit distributions to detect fraud is not at all unreasonable.
http://en.wikipedia.org/wiki/Benford%27s_law
Here's an article that concludes that the election results do not run afoul of Benford's Law:
http://www.jgc.org/blog/2009/06/benfords-law-and-iranian-ele...
I'm calling bullshit.
You gave no reason to be skeptical, and simply illustrated your lack of familiarity with basic tools of probability. Funny that you would below speak of a "stunning example [...] of mathematical ignorance." If you wanted to substantively contribute to the discussion, perhaps you should try explaining why this was an invalid test instead of just labeling the authors as "numerology detectives?"
Your post amounts to "I don't understand statistics or what the hell this article is talking about, therefore I'm going to be suspicious of it."
What would be mathematically sound would be to agree on a standard set of tests, perhaps including Benford's law, digit distribution, two digit combinations shown to be popular when humans pick numbers at random. And to do so _BEFORE THE DATA HAVE BEEN AVAILABLE_.
How would I trick people into believing a false conclusion? One way would be to find 30 or so measures of what could suggest suspicious patterns. Then I'd test the data with them, and perhaps there would be 2 which have a 5% chance of occurring purely by chance. Then I say "Aha! Look at this! The chance of these two things happening are .05 * 0.05 or practically impossible!"
Apparently I could fool some pretty smart people this way, since they obviously aren't immune to basic statistical fallacy.
As another poster said, none of this suggests that the data weren't manipulated, but just that the stats in the article, as given, are purely bullshit.
It could turn out that if the set of 30 tests are all "good tests", then the chance of your being able to find 2 tests that report low numbers like that approaches 0%, in which case you would simply be wrong.
I am not a statistician, and statistics is notoriously difficult for humans to deal with [1], so until I find an answer to the question above I'm going to maintain a neutral position in this argument. Unless you happen to have a doctorate in statistics, in which case I will take your word for it. ;-)
[1] TED talk: http://www.ted.com/talks/peter_donnelly_shows_how_stats_fool...
(30 choose 0)*(.95^30) + (30 choose 1)*(.95^29)*(.05^1)
= 0.553542075
So there's about a 50-50 chance (obviously I'm assuming that the tests are independent of each other.)The relevant Wikipedia article:
The question was: What is the chance of being able to find exactly 2 tests?
Sure thing. That's:
(30 choose 2) * (.95^28) * (.05^2)
= 0.258636738
I thought you meant "at least two tests", for obvious reasons. (If 4 tests showed vote rigging, the researchers would still report vote rigging...)> (and does that mean less than or equal to 5%?)
Yes.
> I thought you meant "at least two tests", for obvious reasons.
Yes, sorry, I did, I guess you meant to say I should subtract 0.553542075 from 1 to get my answer.
(just saw earl's post above...)
Second, what you have calculated is... I don't know. You are calculating the cumulative distribution function for a binomial distribution with n=30, p=0.95. How on earth does that relate to the previous post? Are you somehow confusing a p-value -- which is a statement about the minimal alpha for a rejection region such that given a set of data X and a test, we would reject our null hypothesis H_0 -- and a probability? Because they have almost nothing to do with each other.
Because if the tests are independent, then something is well and truly broken! Your bernoulli CDF is for independent draws.
Further, consider a hypothesis test. You feed it a desired chance of type I error -- say 0.05 -- and it gives you rejection regions for your test statistic. That is, it picks some subset of the domain of the test statistic and labels that as "reject H0" and labels the complement as "accept H0." Claiming that these 30 acceptance / rejection regions will somehow be independent of each other isn't true -- since our tests are accurate, or at least reasonably so, the accept / reject regions will be the same, or nearly so.
Could it be possible to find tests with totally conflicting accept / reject regions? I'd say that's highly unlikely -- the regions would then be complements of each other. What is possible is that you will find paired tests with AR regions that are slightly different. But now you are asking the probability that your test statistic lands in conflicting areas of at least 2 pairs of these test regions. Again, this can be modeled -- try MCMC, perhaps -- but it will not be anything like bernoulli.
Further,
Does this make any sense?
Ok, here's an example. Test 1 tests the following hypothesis: the last two digits of polling data are uniformly distributed amongst the values 00, 01, ..., 99. Test 2 tests the following hypothesis: the first digit is distributed according to Benford's law. For large data (like vote totals) these tests are (nearly) independent, no?
Grab their data: Iran_2009.csv Then: (apparently I can't embed preformatted text. Sorry.) Try this: http://img.skitch.com/20090621-xrmw8yrxaec15jku414ffb2cjq.pn... -or- http://earlh.com/HN/iran.R , and http://earlh.com/HN/Iran_2009.csv
Here's a histogram: http://img.skitch.com/20090621-ewdbtr9q2cw5hyemn64fg3kbqd.pn...
The p-value I get is p-value = 0.07685
NB: there is something weird going on with their data: it doesn't sum properly. http://img.skitch.com/20090621-dx6nqws4u9rc5p3nghsnuy24j3.pn...
--------- So, the way this works is that I look at the last digit in each of the totals, then ask what is the probability that we should observe such a distribution of final digits if the final digit were distributed as discrete uniform U[0,9]. I use a chi2 goodness of fit test, then calculate the p-value with H0 = the final digits are distributed uniformly. We would reject H0 and say the election was rigged at an alpha level of 0.9, but not at 0.95. Take of this what you will, but I'd hope to have a higher significance level for elections. : shrug :
Code for above:
d <- read.csv(file='~/stuff/earlh/Iran_2009.csv', header=T, sep=',')
lastDigit <- function(v){
v - 10*floor(v/10)
}
digits <- lastDigit( c(d$Ahmadinejad, d$Karroubi, d$Mousavi, d$Rezaee))
hist(digits, breaks=10)
#chi2 gof
tab <- table(digits)
n <- length(digits)
model <- chisq.test(x=tab, p=rep(0.1, 10))
model
# hand generated -- check our work above
ts <- 0
for(i in 1:length(tab)){
ts <- ts + ( tab[[i]] - 0.1*n)^2 / (0.1*n)
}
qchisq(p=1-0.076, df=9)Edit: Very surprised at the -3 mod. Are people here really so bad at statistics that they think this?
Example: Just because the average rainfall in a spot is 10 inches, does NOT mean it has to be 10 inches every year. If you have 6 inches one year that is NOT an anomaly. That is perfectly normal. Sure it might be unlikely - but unlikely things still happen! In fact they must happen, just not as often at the more likely things.
So the analysis of the votes, sure it might not match the average, but that doesn't mean a whole lot. It could simply mean that you happened to get the 10% chance.
PS. Not that the vote fraud is in question. It's not, it's obviously fake. But not because of this analysis.
Either if the elections are rigged or not, 49% of the population can not be subdued to the will of the ruling 51%
We need to change electoral systems, specially "winner takes all" ideologies, the tyranny of the majority.
We need to learn to live together and share power together.
And to separate if we can not. Nothing should be imposed.
In Iran, trust in the system was lost. This is much about Iran's supreme leader as it is about the election for president.
You are absolutely correct, the problem in Iran is not a lack of democracy, it is tyranny through democratic means. Just because you win an election, does that mean you should be able to do something that the other 49% of the people find intolerable? It depends on whether you believe in democracy as a guiding principle, or in democracy as the letter of the law. The belief in democracy as the letter of the law, is, in my opinion, democracy taken to the extreme. It is the definition proffered by those who would use democracy as a cover for wanting, winning, and wielding power over those who may, legitimately, disagree with them.
Democracy is a beautiful idea. Now, thanks to the efforts of politicians in democracies of all stripes, not just those in Iran, it has become the principle endorser of divisiveness in the world of ideas. It is being used to give violently divisive ideas an air of legitimacy that they do not deserve.
This system has the advantage that minority views are not only heard but also have votes. And since politics is a game of agreements, coalitions and I'll scratch your back if you scratch mine even small parties have a chance of getting legislation passed if they are willing to compromise on something else. It equates to more consensus seeking legislation, and an often vocal minority that actually has political influence.