The real problem was a lack of high-quality state-level polls, and in one or two cases major misses on what state polls there were; there wasn't all that much visibility on many states.
And ultimately, the final result was so close that it was below the margin of error of any reasonable poll to detect. Even assuming the highest quality polls conducted in each state, every day, the best they'd have been able to say would have been, pretty much "It's 50/50", in retrospect.
That wasn't even really a problem; the error in state polling was pretty consistent with the error in national polling; the problem with predictors other than 538 is they treated state variations from polling averages as independent, 538 correctly assessed them (based on past evidence) as highly correlated, which is why Trump had a nearly 1 in 3 chance in 538s forecast.
The problem is people taking (honestly or just for the purpose of after-the-fact criticism) "1 in 3" to mean "absolutely won't happen".
That would be hitting on a 30% chance in 2016 and hitting on a 20% chance in 2020. That's a 6% chance overall.
6% likely outcomes happen all the time. It would be a textbook example of "resulting" to draw a conclusion based on this.
So while you say 6% "happens all the time" - no it actually happens 6% of the time. But 40% for both terms would indicate more like he had a 70% last time and maybe another 60% this time. The pollsters can be correct for fucking California, but what good is that if they're super wrong because racist people in Alabama don't click or pick up the phone (or have no phone to begin with) or talk to pollsters.
To clarify, my intention was to state, "outcomes with 6% of odds happening will occur very frequently" - not relatively. A good hitter in baseball hits homers about 6% of the time. If they hit homers in 2 at bats, we would not have enough information to say much about their true talent.
That's true only if you assume perfectly spherical pollsters :) After a real or perceived polling miss (in reality, most polls weren't off by all that much in 2016), pollsters tend to change assumptions, sometimes overcompensating. Notably, after a moderate polling miss in the 2015 elections, which predicted a hung parliament where in fact the Tories managed a small majority, UK pollsters went on to overcompensate in the 2017 election, showing a Tory blow-out win whereas in fact they ended up with a hung parliament.
(That one, incidentally, shows one danger of polls! The first poll allowed Cameron to safely promise a vote on the EU, because he assumed they'd be either out of government or in coalition with the libdems, who wouldn't allow it. In fact, they won and were forced to go through with it. The second one allowed Theresa May to call a snap election to gain seats so as to have enough spare MPs to be able to ignore the ERG and negotiate a semi-sensible Brexit. Instead, they lost seats, had to deal with the ERG and Unionists, and ultimately ended up with the current catastrophe. The current Brexit mess is at least in part due to polling misses.)
In the US, pollsters have started to pay a lot more attention to education, which was a major predictor in the 2016 election.
I'm starting to believe that sites like 538 feed off of the general public not understanding that a 20% chance to win is 1 in 5. Endless ink has been spilled about how the public does not understand statistics but very little effort has been made to communicate effectively.
The whole thing seems like a navelgazing sideshow where the pollsters want to have it both ways. They want to claim their predictions are infallible but then when the public says they failed they want to backpedal with holier than thou "well, actually..." excuses.
How, then, should statistics be appropriately communicated?
> They want to claim their predictions are infallible but then when the public says they failed they want to backpedal with holier than thou "well, actually..." excuses.
Do pollsters generally try to claim their numbers are infallible? If anything, the fact that margins of error are included in the results would seem to imply the opposite.
*https://pbs.twimg.com/media/Cwc-j6YXUAEIMaM.jpg
The pollsters take too much credit when they're right and deflect too much blame when they're wrong. They need to do a better job of being humble and stop trying to pass themselves off as apolitical number crunchers just giving us the facts.
You are confusing pollsters with "forecasters using data from pollsters". These are not the same people.
Many of the forecasters in 2016 were very bad because of naive assumptions about how polls related to election results, particularly many making the assumption that polling errors were independent between states (also, lots did a really poor job of poll aggregation before those naive models.) The reason is Nate Silver and 538 had made a name for themselves in the preceding couple of cycles, and lots of people who didn't understand the process but saw the outcome decided they could do the same thing, because, hey, how hard could it be to compile a bunch of polls and model outcomes based on them?
In particular look at this (which is almost exactly what happened):
But if there’s a 3-point error against Clinton? That would still leave her with a narrow lead over Trump in the popular vote — by about the margin by which Gore beat Bush in 2000. But New Hampshire, which is currently the tipping-point state, would be exactly tied. Meanwhile, Clinton’s projected margin in Michigan, Pennsylvania and Colorado would shrink to about 1 percentage point, while Trump would be about 2 points ahead in Florida and North Carolina. It’s certainly not impossible that Clinton could win under those circumstances — her turnout operation might come in really handy — but she doesn’t have the Electoral College advantage that Obama did in 2012, when he led in states such as Ohio and Iowa and had larger leads than Clinton does in Michigan and Pennsylvania. In particular, Clinton could be vulnerable to a slump in African-American turnout.
What would "good faith" communication look like, then?
> I remember very clearly sources like this* in 2016.
Can you elaborate on what exactly is wrong with that source?
(That isn't a pollster, too; that's a (meta?)analysis/aggregation, but that's a relatively minor nitpick)
> They need to do a better job of being humble and stop trying to pass themselves off as apolitical number crunchers just giving us the facts.
Are they themselves claiming that they are "giving us the facts", or is that how they are being represented by others?
And are you talking about actual polls, or about predictions based on said polls?
It does a pretty good job - it gave Trump nearly 30% chance of winning in 2016 and given his small margin that seems reasonable.
Sites like Huffington Post which gave Clinton 99%+ chance are the ones which should be criticized.