Integer percentages as fingerprints of electoral falsification
arxiv.org
arxiv.org
I wonder if someone has some free time to run the models on several recent elections that were filled with fraud. I'm interested in: Austrian presidential elections [0], Turkish elections [1] and Democratic party primaries [2][3]. Each of them was analyzed for voting fraud and fraud was found in each of them (some were full of massive vote fraud like Turkish elections).
PS: What's with all the downvotes? Is the evidence of massive voter fraud in the West so unsettling that you have to downvote me?
[0] >Austria presidential poll result overturned
http://www.bbc.com/news/world-europe-36681475
[1] >Turkey Elections Massive Vote Fraud
https://erikmeyersson.com/2015/11/04/digit-tests-and-the-pec...
[2] >Hillary Clinton Favored By Election Fraud In Democratic Primaries
http://www.inquisitr.com/3127046/hillary-clinton-favored-by-...
[3] >Election Fraud Watch 2016 (J: awesome blog that tracks election fraud in primaries)
As if anyone should be surprised that a leftie social democrat wouldn't do well in a US election.
I don't care if I get down voted into oblivion for saying this next part: When looking at the spread between the types of polls, there is a massive difference between the voting elite and the people's will. I don't advocate against the electoral college, but what was supposed to be system that prevented buyouts has proven to be completely forsale by the Clintons. I didn't see massive pro-Clinton rallies. We had quicker investigations into Samsung VS Apple for inconsequential things. No one has been held accountable and no explanations have been given for massive breaches of protocol other than its a plot by the Repuglicans.
I'm tired of everything from the Clinton family. My state KS has a good chance to throw electoral votes at Gary Johnson. Please California if you are listening, don't reward this family for their actions.
(I can't stand the Clintons, but I don't think they cheated)
- The recent release of hacked documents from the DNC show that the primaries were not set up to be a fair competition from the start. - The rules are specifically set up in a number of states (New York is a prime example) to disenfranchise any new Democratic Party members from being able to vote. - Thousands of registered voters were removed from the voters roll. There is even a case where a lifelong Democrat running for a congressional seat on a Democratic ticket was removed. - Media coverage made it appear that Clinton was in the lead before any voting even took place. That coverage has been shown to have been orchestrated (at least at the start) by the DNC. - Funds raised in the DNC's name were disproportionately funnelled to the Clinton campaign. - Voting stations were closed down in huge numbers in some states and territories (Puerto Rico and California being prime a examples)
It's going to look pretty ridiculous if the DNC is ever audited, or has someone more credible than Guccifer2.0 hack and release their documents.
I.e. 50% of sanders' voters were women. You'll get similar proportional percentages of minorities.
Simpson's paradox creates the perception that sanders' voters are composed of mostly white males. It's untrue.
Some people try to explain it away by saying people vote for one candidate but are ashamed to admit it to a person asking them who they voted for, but wouldn't that tendency show up statistically across all precincts rather than in specific ones?
There is also an alarming number of times this occurs specifically in precincts with electronic voting machines.
ashamed to admit it
Eh, no need for shame, it could just be that one candidate appeals more to busy (or privacy-minded) voters who decline to be polled.I generally avoid people with clipboards who are approaching people on the street, and I can believe that's something that correlates with opinions on other things.
Not downvoting you. But if fraud turns out to be the case, I'm going to be unsettled as hell. It should be unsettling if you live and/or care about democracy in the West, obviously. I doubt we disagree about this though :)
But even if no fraud is suspected or if it finds nothing, this is data science. Unlike say medicine or food science, there's no shortage of data (in theory). So if there's a country that we can reasonably be sure about their democracy is entirely above the table[0], throw it into the test and use that to calibrate the others.
One problem I believe is that there is no such thing as "the test" :-)
This research was done based in a particular system of elections--if you look beyond your own democratic system (or the way you've learned in school how it's supposed to work), you'll know how insanely different they all are. The fact that we actually use a single word "democracy" to describe them all is pretty crazy, now that I think about it.
Both the numbers themselves (what aggregate counts are available) and the specific way of fraud suspected (which you kind of need for a statistical hypothesis) depend heavily on the democratic system involved. For instance that messy thing with the US elections a while back, second term of Bush IIRC, whether there was fraud or not, it was fought over a couple of hundreds of votes in one district or county, was it not?
To disappear or push a few hundred votes one way or another requires a different type of fraud, and therefore statistical testing, than say, the type of fraud that makes a whole country of several hundred million almost always divide their votes between two options so close to 50/50 that a few hundred votes can swing the outcome in the first place.
And that's just one country. So much different creative ways to mess up whatever is locally deemed to constitute democracy, you probably need at least an equal amount of different creative ways to statistically test whether they're fraudulent or not.
[0] come to think of it, maybe there is a shortage of data ...
Does anyone know why Benford's law doesn't work here but does work for other made up numbers in applications like accounting?
This isn't exactly new, such analysis was done as early as 2011, i.e. http://lleo.me/dnevnik/2011/12/07_gauss.html
Although, they were more focused on the peculiar distribution, than on integer spikes.
I'm tempted to argue the improbably-round numbers might be due to lazily counting/sampling the ballots rather than actual malicious fraud... but I guess sloppily running the election still constitutes fraud in some sense.
So, while theoretically possible, realistically summed up by "three men can hold a secret, if one of them is dead."
(And specifically here, if I were a conspirator going for a predetermined result, I would want a somewhat close result, yes - but not an almost-tie as seen here. Something like 55:45 or somesuch, not a "every single vote matters" scenario, where a few thousand votes could swing the outcome)
There are other problems associated with round ballot counts. Here is a funny one: a sizable number of polling stations reported round number of counted ballots, coinciding exactly with the total number of ballots received by the polling station prior to the election day (this information is available too). As if so many people came to vote that every single ballot was used up, turnout being close to 100%. There is little doubt that it was all made up out of thin air.
Recently on HN there was a related test (the Grim test) https://news.ycombinator.com/item?id=11787560
Russian elections for statistics have become like "Lena" for image processing. Typical "camel-double-hump-Gauss" of Russian elections - polling stations without observers at Fig 2. at http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3545790/pdf/pnas... ("United Russia" is the Putin's party).
I guess this means they found the data first, then decided to try to analyze it and compare it with similar data elsewhere. It's probably not their work-related research (http://www.fchampalimaud.org/en/the-foundation/mission/).
"If you tell the truth, you don't have to remember anything." - Mark Twain
The election results in Crimea and Donbass do contain some funny patterns and are most likely largely made up, but no information is available on the polling station level so there is little to analyze statistically.
>> no information is available on the polling station level so there is little to analyze statistically.
What they did was look at the percentages and exact vote counts for each choice, and noticed that they match down to one vote (i.e. if you take the percentage, multiply by total population, and round, you'll get the vote count). For example, in referendum in Lugansk, the official stats are as follows:
1,349,360 valid bulletins 1,298,084 (96.2%) voted yes
Normally, the percentage is a number with a lot of digits after the decimal point, which is then rounded for presentation purposes. But here, if you take 96.2% and multiply it by 1,349,360, you get 1,298,084.32. In other words, it's actually accurate to 4 digits after decimal point, three of which happen to be zeroes (96.2000%).
Of course, it could be just a very unlikely (< 1/1000) coincidence that the number of "yes" votes just happened to fit exactly into three digits of precision. But a more reasonable explanation is that someone started with the percentage that they wanted to get, and then computed the requisite number of bulletins from that.
Said explanation becomes even more reasonable when you take the reported numbers for turnout, and realize that the same relation holds there. Specifically:
1,807,739 eligible voters 1,359,419 voted (75.2%)
1,807,739 х 75.2% = 1,359,419.73
Oh look, another perfect percentage (75.2000%).
Likewise, I'd expect there to be proof of fraud by other reliable means in order to validate this method. It is not enough for them to just assert that there can be no other explanation for this data, so these were fraudulent results, so our method must be working.
Absent a control, the strong conclusion that fraud can be detected this way seems unsupported.
Would you kindly tell me the method they used to independently prove fraud in some of the analysed data.
Having now read the article, I stand by my original point. There seems to be no verification that the proposed method does indeed identify (in some cases) cases of elections that were independently known, as a point of fact, to have been fraudulent.
They seem to be claiming two things, but you cannot have it both ways :
a) This method can identify fraudulent elections and we proved it by analysing elections in Russia
b) These elections in Russia that we analysed using this method were fraudulent - as proved by the method.
I would expect to see a control that takes data from elections that were already known to be fraudulent. e.g. by confession, video evidence or some other reliable means to show that the effect was observable in some of those demonstrably fraudulent elections.
Not "as proved by the method", but "as proved by the science of statistics". So let me fix it for you:
a) This method can identify fraudulent elections, because it's scientifically correct.
b) These elections in Russia that we analysed using this method were fraudulent - as proved by the science of statistics.
FYI, the paper was published in "Annals of Applied Statistics", which is a peer-reviewed journal.
I have asked this question in different ways, but to my (admittedly limited) understanding, the question has not been answered by you or anyone else here.
Surely there would have been a better set of test data than one for which politicians are still actively serving.
Unfortunately the science is somewhat tainted by the politics here since it is very unlikely they'll get independent confirmation of fraud for the test set.