Was Nate Silver the Most Accurate 2012 Election Pundit?
appliedrationality.org
appliedrationality.org
What's a fair benchmark? This article offers up a "coin flip" for each state, computing that such a coin flip would have a Brier score of 0.25. (The Brier score is a mean-squared error between outcome (1 or 0) and the percent certainty of the prediction in that outcome. If a coin flip is the model, each state's result of 1 or 0 would be in error by 0.5. The mean squared error would be 1/51 * 0.25 * 51 = 0.25.)
But... that seems like too generous a benchmark. Take the simple model: "assume 100% likelihood that state X will vote for the same party as it did in 2008." That guarantees that deeply red or blue states will vote the same way, so it takes the non-battlegrounds out of the equation.
With this model, there would only have been 2/51 errors. This simple lazy model achieves a Brier score of 0.039, beating Intrade and the poll average computed in this article quite badly.
After working through this, I'm still impressed by Silver and the other quant predictions. But I'm more concerned about media that rely too much on reporting a single polls result as "news" rather than as part of a larger tapestry.
Then again, it's the maligned media polls that are the raw input to Silver and the other models. Unless the media keeps funding the polls, the quality of these more advanced models will suffer.
(I don't know when the new numbers will go live on the blog; Luke handles that.)
Is there actually evidence that higher poll numbers in favor of X lead to higher voter turnout from supporters of X? It seems like everyone takes that for granted but I've never seen any evidence that it's true.
I'd like to see some raw evidence, too.
But since we're forced to speculate at the moment, my bet would be that the relation to voter turnout and a candidate's chances of winning isn't linear. It's probably more of a parabola: the more extreme a candidate's chances of being elected (or being defeated), the less likely anyone is to come out and vote, because it feels impossible to make a difference. On the flip side, I'd suspect that the closer and more contested a race is, the more likely people are to feel obligated to vote.
Under this theory, polls/predictions can't actually skew the result in either direction [1], they can only increase or decrease the total turnout (highest turnout when polls say 50%, lowest turnout when they say 0%/100%). In reality this probably isn't actually true, but we can't say for sure whether or not polls do affect election outcomes, and in which direction, unless there's empirical evidence. I don't even have a guess for which direction it would go (whether you want your supporters to be "concerned" or "optimistic") - I could see either being true.
[1] Edit: I should note that this assumes that each poll/prediction is listened to and taken seriously equally by both sides of the electorate, which almost certainly isn't true. Which actually brings up something kind of interesting - it could be that with all of its "Romney will win in a landslide" talk, Fox News actually hurt Romney's chances because most Fox News viewers are Romney supporters (but they could have done the same amount of damage to Romney's chances by saying "Obama will win in a landslide" - the best way for them to help Romney might be to say that the election is exactly tied). And I suppose it would also lead to a justification for the idea that Nate Silver has a liberal bias, using his relatively "wishy-washy" predictions to energize his mostly young and liberal audience to get out and vote. It seems that I've gone too far with this and started arguing against myself...
However, in a winner takes all election, the effect would have to be huge. Let's say all polls indicate a 51%-49% result. Then, at least 4% of the winning party's voters would have to stay home 'because they already won' to change the result (and that assumes none of the other voters stay home 'because they already lost'). At a more realistic 60%-40% poll prediction, one in three voters would have to stay home.
I have no evidence for this whatsoever.
I think my point still stands though: this doesn't showcase NPR's neutrality - perhaps they are in fact biased toward Obama and just lie more strategically than most Republican pundits. (Certainly not accusing them of that, only saying that I don't think we can glean much from the fact that they gave Romney higher numbers than he deserves.)
I'm not sure it's really a question of ideological bias so much as filtering the raw data through a backward-looking model.
The former can occur simply because of the assumptions that your model makes. Bias in "statistical bias" does not mean the same thing as "political bias".
(I used to be part of that research group but I left at the end of 2010 to start Beeminder.)
Actually, that makes this whole exercise premature, the rankings may change once the final data is in.
Another thing that looks odd on that graph: the given polling numbers from Washington Times/Politico/Monmouth/Newsmax/Gravis/Fox/CNN/ARG all look identical despite their differing margins of error (which suggests their source data is different). What's going on there?
"According to US census data, just 71% of eligible Americans are registered to vote. In 2008, almost 90% of those who were registered did vote. So in any poll, it is vital to know which respondents are on the register."
But there're a few more interesting subtleties in there too, so it's worth a read.
I dunno. IIRC, Drew Linzer of Votomatic even worried on his blog that the polling numbers were too close and that pollers might be fudging their numbers to be more similar to each other (which would lead to substantial overconfidence in estimates). Still, the final results seem pretty accurate, so...
Except this seems to ignore that P(comment about Sam Wang | I am Sam Wang) is much higher than P(comment about Sam Wang | I am not Sam Wang) :-)
Edit: The confusion is probably that Silver was reporting the average electoral split as his prediction, when the mode is more important in what you're talking about. His average was almost never a number that was actually possible, since he was quoting them to the nearest 1/10th of a vote, so its kind of unfair to punish him for not getting it exactly right.
http://www.slate.com/articles/news_and_politics/politics/201...
Silver takes all the polls, even the crappy ones, and includes them in his calculation. If you're a bayesian you'd find this comforting because all of the evidence is included in the belief. There might be some handwringing about how important each poll really is. If you're not a bayesian, then you have some other weird strategy that might or might not work.
So really, the predictions are nice, but what we're after is a system that produces good predictions. It's not clear that Mr. Silver's is the best. Perhaps there is some horrible flaw an evil agent could exploit that just didn't get tickled this election. It's tough to say.
Off-topic, but this is my beef with the "Bayesians" - it's not all of the evidence; it's completely absurd to believe that all evidence can ever be accounted for, when considering anything.