The fivethirtyeight R package
blog.revolutionanalytics.com
blog.revolutionanalytics.com
I'm getting really irritated with people (in this thread and other) attacking 538 along partisan or personal lines for not making the same predictions they did personally. 538 made a model, not a poll; they are only ever as good as the polls they represent, and even then, they still managed to be better than quite a lot of public polling and commentary. If Hillary Clinton had won, the same kinds of people on the other side would be attacking 538 for being too bullish on Trump.
The worst part is that a lot of this is predicated on the idea that giving something a 1 in 3 chance is the same as saying it'll never happen, when it's almost the opposite. There are some critical misunderstandings of statistics going on here, on all sides.
On the morning of the election, FiveThirtyEight said that the chances of Trump winning were roughly equal to the chances of a random coin flip coming up heads twice in a row (25%).
In that situation, a good statistician would certainly say that the expected outcome is for at least one of the coin flips to land on "tails" (75% chance). But if the coin lands on heads twice, you wouldn't say that "statistics is broken"[0] or that the statistician predicting the outcome was wrong.
Similarly, let's say that you had a separate person predicting the probability of the double-heads coin flip to be 90%. If it does happen, that doesn't mean they made a good prediction, just because the outcome was correct. They predicted the correct outcome, but the degree of confidence they expressed in it was completely wrong.
Predictions aren't just about getting the outcome right - they're also about expressing the right degree of confidence in the results.
[0] Which is literally what some otherwise-respectable journalists complained after the election
(Of course, the degree to which the data was lacking and the reasons for that lack is another issue in and of itself.)
In fact, that very thing did happen in the week leading up to the election, so sure were some that Clinton would win. There was a raft of articles nitpicking the model, particularly the adjustments that it applied to polls and the level of uncertainty that it assumed.
Ugh, yes, like that HuffPo guy attacking Nate Silver. He apologized later, at least.
Systemic polling error is a thing that happens. It's happened in previous elections. The polls are much more accurate than chance, but they aren't infallible. Particularly in this election, it's difficult to predict voter turnout from polls.
538 was one of the only models that took that into account. And as far as I know, they gave better odds to Trump than everyone else that tried to predict the election with statistical methods. Certainly they have a better track record than political pundits, which have never been better than chance.
To the extent 2018 brings increasingly contentious elections, we will likely see polling to remain inaccurate due to voter unwillingness to state unpopular opinions. In that case, multi-input "big data" approaches may prevail.
(Of course, it's also possible that 2016 is a "top" in terms of divisive rhetoric...only time can tell.)
What do you think this would look like? Inferring voter preferences based on proxies or instruments?
plus do polls take into account where they are polling, some areas are hostile to one party or another and they may not get a good response regardless of intent
This was intended to pick up on the Shy Trump voters.
Granted hindsight etc.
I'd say the same will be true for some of the middle-right wing parties in Europe. Notably I think AfD in Germany will (unfortunately) get more votes than polls will indicate in the upcoming election(s). Le Pen in France could be a similar case but I'm not well informed enough on French politics to feel sure about that statement.
Even if the model said 95% for Clinton, there's still a small possibility for Trump to win. It's hard to accept because 95% is really high! But hey, people win the lottery, right?
This is a rare case where the "data-driven" cult is forced by reality to realize how worthless everything it praises actually is, because data doesn't measure the right thing. Normally people simply ignore the difference between what's measured and what you think you're measuring and all the wrong data-based conclusions become religion every civilized person is expected to believe in. That's why it's important to drive home the point how badly 538 did (and yes, others did even worse, but 538 didn't do "pretty well", it did tremendously badly.)
And BTW there were better "data-driven" predictions than 538's, just not based on polls. Allan Lichtman's model is one that got it right. Incidentally, the incentive to outright falsify polling data (not only for pollsters but to some degree, for those polled!) is much larger than the incentive or the ability to falsify the kind of data his model looks at.
Finally, someone who worked at 538: "It sometimes seemed as though 538's interpretation of the math wasn’t free from subjective bias" (https://www.theguardian.com/commentisfree/2016/nov/09/polls-...)
June 16, 2015: Why Donald Trump Isn’t A Real Candidate, In One Chart
July 16, 2015: Two Good Reasons Not To Take The Donald Trump ‘Surge’ Seriously
July 20, 2015: Donald Trump Is The World’s Greatest Troll
Aug. 6, 2015: Donald Trump’s Six Stages of Doom
Aug. 11, 2015: Donald Trump Is Winning The Polls, And Losing The Nomination
Nov. 23, 2015: Dear Media, Stop Freaking Out About Donald Trump’s Polls
Donald Trump Comes Out Of Iowa Looking Like Pat Buchanan 8 Nov: There's A Wide Range Of Outcomes, And Most of Them Come Up Clinton
6 Nov: How Much Did Comey Hurt Clinton's Chances?
6 Nov: Don't Ignore the Polls – Clinton Leads, But It's A Close Race
5 Nov: Why We Don't Know How Much Sexism Is Hurting Clinton's Campaign
4 Nov: National Polls Show Clinton's Lead Stabilizing – State Polls, Not So Much
4 Nov: Trump Is Just A Normal Polling Error Behind Clinton
3 Nov: Why Clinton's Position Is Worse Than Obama's
2 Nov: The How-Full-Is-This-Glass Edition
2 Nov: Trump's Chance Of Victory Has Doubled In The Past Two Weeks
1 Nov: Yes, Donald Trump Has A Path To Victory
1 Nov: On A Scale From 1 To 10, How Much Should Democrats Panic?
31 Oct: The Odds Of A Popular Vote-Electoral College Split Are Increasing
31 Oct: Comey Or Not, Trump Continues To Narrow Gap With Clinton
27 Oct: Don't Read Too Much Into Early Voting
26 Oct: Is The Presidential Race Tightening?
24 Oct: Why Our Models Are Much More Bullish Than Others On Trump
You get the idea.Trump has as much chance of winning as Cubs
sli.mg is being attacked so I can only post a link you can't see !!
He was described as an "insider pundit" by the Clinton team
https://wikileaks.org/podesta-emails/emailid/28878
From:brentbbi@webtv.net : ... I see Dan, Nate Silver and various other insider pundits saying what they are saying, I have been to enough rodeos to know that thoughts are being planted from the Obama-Clinton consultant class (not you, and you know who I mean). This is not helpful to Hillary, quite the contrary, ...
So I don't consider him objective.
> Trump has as much chance of winning as Cubs
Are you referring to this “The Cubs Have A Smaller Chance Of Winning Than Trump Does”[0]? If so, I don’t understand why you would consider such a headline to be inaccurate or not objective.
[0] https://fivethirtyeight.com/features/the-cubs-have-a-smaller...
September 15, 2016: How Trump Could Win The White House While Losing The Popular Vote
link: http://fivethirtyeight.com/features/how-trump-could-win-the-...
All those people need to do is go to the election forecast page, which is still up[1], and look at the percentages under "Popular vote". Or down lower under "How the forecast has changed", click POPULAR VOTE to see the percentages over time.
For instance, the final pre-election estimate before the election gives:
- CHANCE TO WIN of 71.4% Clinton vs. 28.6% Trump, around a 1 in 3.5 chance for Trump to win.
- RESULTS ESTIMATE of 48.5% Clinton vs. 44.9% Trump, with a big overlapping margin of error.
[1]https://projects.fivethirtyeight.com/2016-election-forecast/
I completely agree with almost all of your post, but want to defend the polls, even though you didn't explicitly blame the polls. FiveThirtyEight themselves consistently defend the polls even in the 2016 election. For example, they ran an article a few days before the election that Trump was only "a normal polling error" away from Clinton [1], and in their podcasts and articles since the election, have stated that in fact an average polling error was exactly what happened. The problem with a normal polling error in this case was that the election was close, and in particular, states that were critical for Clinton were very close. And FiveThirtyEight's model was the only one to really recognize the degree of uncertainty this really created by giving Trump a 25 or 33% chance.
[1] https://fivethirtyeight.com/features/trump-is-just-a-normal-...
The vignettes (https://mran.microsoft.com/web/packages/fivethirtyeight/vign...) are a good tutorial in dplyr/ggplot2 too.
https://www.r-bloggers.com/interactive-visualizations-with-r...
None of them are as flexible as D3, but there are packages to output from R to D3 out there.
Here's a higher-level demo of ggplotly in one of my R Notebooks: http://minimaxir.com/notebooks/breach-network/
Shiny requires an external server and is often overkill for static data.
(You can build a Plotly chart in ggplot, embed it in a web page and then script it with JS to get a visualization. Frankly, that's a PITA. You can write R/shiny code that will generate a visualization, with HTML controls etc, and you end up writing a web app in R and that's a different PITA. I would like to see Plotly generate a web wrapper that automates wiring up a plotly graph and scripting it from an app so you can update data on the fly to make it a visualization. Supposedly they are working on it with their dash framework, but it's not released yet.)
The embed-directly-into-HTML is the approach I use for my Jekyll blog and it does not require much effort. (I just have to set a YAML flag to load Plotly library)
The part that requires some custom Plotly coding is when you want to script the chart ... have HTML buttons, sliders, to filter, recompute, revisualize the data.
It's beautiful that it even works, and you can programmatically update the JSON of an embedded chart. But it would be nice if you could give Plotly a dataset, plot the data, and say, give me a button to filter rows using these criteria for this column (or run some other code on the data) and re-render the chart. Right now you have to code that manually and update the JSON model.
Visualization used to mean charts that were animated and/or would dynamically update but now it means any old chart LOL.
There's some interesting datasets in here[1]! Everything from "How American's like their Steak"[2] to "The Most Common Unisex Names In America: Is Yours One Of Them?"[3]
As a (mostly) Python person I'd like to see these in Python!
[1] https://mran.microsoft.com/web/packages/fivethirtyeight/five...
[2] https://fivethirtyeight.com/datalab/how-americans-like-their...
[3] https://fivethirtyeight.com/features/there-are-922-unisex-na...
Here is their methodology -- apparently quite flawed?
http://www.huffingtonpost.com/entry/high-probability-clinton...