Why we should care about the Nate Silver vs. Nassim Taleb Twitter war
towardsdatascience.com
towardsdatascience.com
As just one example, the whole digression on a “decision boundary” is conceptually mistaken. Once you report a posterior probability, it’s up to the user to establish a simple real number threshold for placing a bet (if we’re talking about means to put skin in the game). The posterior you just learned holds all information, your only action is to threshold it. That’s the Neyman-Pearson lemma.
If you allow generalizations to interval-valued probability, which might be sensible under these conditions, the situation gets more complicated. But the writer of the post did not mention this.
Another place this came out is his initial discussion of aleatory vs. epistemic uncertainty — this can be done clearly, but here we read:
> Aleatory uncertainty is concerned with the fundamental system (probability of rolling a six on a standard die). Epistemic uncertainty is concerned with the uncertainty of the system (how many sides does a die have? And what is the probability of rolling a six?).
This is just not helpful. He has said “probability of rolling a six” twice.
Source: do UQ in day job.
> Because FiveThirtyEight only predicts probabilities, they do not ever take an absolute stand on an outcome: No ‘skin in the game’ as Taleb would say. This is not, however, something their readers follow suit on. In the public eye, they (FiveThirtyEight) are judged on how many events with forecasted probabilities above and below 50% happened or didn’t respectively (in a binary setting).
Saying "this is not something their readers follow suit on" - only if you're a reader who doesn't understand probability. In fact, Silver is constantly trying to get across how likely < 50% probabilities are with his football analogies (e.g. "a team down by 3 at the half actually winning") to emphasize how likely something like "20%" actually is.
If anything, I love Silver's reality-based approach exactly because he does acknowledge that all he can do is offer a best estimate based on available evidence, instead of saying things with false-certainty and then basking in self-proclaimed genius when luck happened to run in his favor.
https://thehill.com/homenews/sunday-talk-shows/414759-nate-s...
> "So in the House we have Democrats with about a 4 in 5 chance of winning," Silver told ABC's "This Week." > "But no one should be surprised if they only win 19 seats and no one should be surprised if they win 51 seats," Silver added. "Those are both extremely possible, based on how accurate polls are in the real world."
Dinesh D'Souza jumps on him for hedging his bets: > And just like that, an 80 percent chance for Democratic takeover of the House goes to 50-50.
Nassim Taleb also jumps on him in the same way: > @DineshDSouza is right: you don't change a forecast from 80% to 50% under uncertainty. Second time klueless Nate Silver makes a mistake. > When someone says are event and its opposite are extremely possible I infer either 1) 50/50 or 2) the predictor is shamelessly hedging, in other words, BS. > 4- Things are actually worse: that @NateSilver538 doesn't get that if BOTH X & Non-X are "extremely possible = Probability converging to 50-50 is EXACTLY my problem w/his misunderstanding of probability in forecasting, & point of my paper. > 5- FOR THE RECORD I am not translating @DineshDSouza's point that it was 50-50: I ALSO understood ~50-50 DIRECTLY from Silver's "extremely" in the linked "The Hill".
The Twitter fight happened on Nov 4, if you want to go look.
In this article, the author notes that 538 only presents its model to readers in terms of probabilities, but the public tends to round their predictions up and down.
> This is not, however, something their readers follow suit on. In the public eye, they (FiveThirtyEight) are judged on how many events with forecasted probabilities above and below 50% happened or didn’t respectively (in a binary setting)
This isn't the public's fault, says the author, it's 538's fault for not clearly presenting the uncertainty in the prediction.
Additionally, the author writes, the 538 polling-based model swung too wildly in the 2016 election. In doing so, it failed to present a single prediction on which it could be judged. It also failed to acknowledge that unforeseen events could change the polls.
Personally, I'd argue that the swinginess of the conveyed the uncertainty of the polls. I spent the last month of 2016 citing Nate Silver to argue on Facebook and Reddit that people were overestimating the certainty of the election forecasts. Fifteen percent of the electorate was up for grabs with just two weeks left in the election! (about 7.5% undecided; and 7.5% saying they'd vote 3rd party)
> This is Taleb’s primary argument; FiveThirtyEight’s predictions do not behave like probabilities that incorporate all uncertainty and should not be passed off as them.
Nate Silver is being blamed both for trying to convey to the public that his predictions involve uncertainty and for the public not perceiving the uncertainties in his predictions.
The only way you could get "50-50" from that is if you read "the Democrats get 51 seats" as "the Democrats win the House", because without that inference, he didn't say anything about the chance that they win the House. He compared a very lopsided win to a slight loss, and those were equally likely.
I'm not arguing with you, since the one thing I know is that there's no surer way to making a mistake than jumping in confidently on probabilistic issues, especially issues of interpretation rather than of pure mathematics; but I think you must have misspoke, or I must have misunderstood. The quoted sentence says 'extremely probable' (an absolute condition), not 'equally likely' (a relative condition):
> "But no one should be surprised if they only win 19 seats and no one should be surprised if they win 51 seats," Silver added. "Those are both extremely possible, based on how accurate polls are in the real world."
Indeed, despite quoting it in my response, I made exactly that mistake. Thanks!
What on earth is a professional like Silver doing using the word "extremely" in that context?
If I'd have written that kind of statement in school, the teacher would have crossed it out in red pen and told me to rephrase.
If I toss a coin, are the outcomes "heads" and "tails" both supposed to be "extremely possible", with "landing on its edge" being merely "unlikely"?
(In case you can't already tell) I'm with Taleb on this one.
Extemporizing while being interviewed on TV, presumably to stress the uncertainty in the prediction, is what he is doing, not writing an HN post, tweet or paper.
I suppose that whenever you are interviewed on national TV, you only say exactly what you mean, nothing more and nothing less?
As for the "50/50" here, I don't think this is meant as being exact numbers (after all, the whole point of the issue is that those exact numbers don't really tell you anything, if anything still is possible and any outcome can be justified later), but simply as the common usage of a phrase in the vernacular for "we don't know either way".
The whole point of probabilistic modeling is to replace absolute decisions like "right or wrong" by continuous weights on the possibilities. If you absolutely need a definite decision, you can sample a prediction according to the probability assigned by the model. If the true outcome is x and the model assigned it probability p, then that procedure is going to be wrong (1-p) of the time. You could define that number as the "wrongness" of the probabilistic model, as a continuous analog of the definite case.
The advantage of probabilistic modeling is that you can also ask how wrong the model expects to be and get a meaningful answer. If there are many possible outcomes and none of them very likely, any choice is going to be wrong a lot. But you should expect a good model to have a small difference between its expected and actual wrongness. One might call that value "honesty".
The whole point about the current topic as well as of my post: You missed by about a thousand miles. Please read it again. It's really pointless to argue about a strawman created by you. Your model is useless, that's the point! It makes no real(!) predictions - not usable for anything apart from blowing ever more hot air, and if it doesn't come to pass, you are never wrong because you left the door open by not actually saying anything in the first place.
Do you not understand that the guy/his company did nothing at all? And that giving some arbitrary probability was/is utterly devoid of any meaning (especially if you can't be wrong whatever the actual outcome)? They could have made any prediction at all, what difference would it have made? That is the value of that "work".
However, I realize there's people who like such meaningless drivel. It is a version of appearing to actually do something while not actually doing anything. You make it into the news but you can never be held accountable because whatever happens happens, you just helped create a few more entirely useless headlines (apart from helping with page views and ad impressions of course). It's actually quite ingenious to misuse actually useful tools like statistics.
I'd never come across anyone using this phrase until Nate Silver did. Is it a US English thing?
It doesn’t mean that a broken heart will definitely kill you, just that there is definitely a chance that it will kill you.
I could totally imagine Nate trying to emphasise the idea that 20% is a much bigger probability than say 1% by using language like that, but put too much weight on it here leading to confusion.
"Possibility" is like "optimality" in the sense that they are binary attributes and thus don't admin grades. Something either is or isn't possible. Qualifying something as "more possible" is a mistake, and qualifying it as "extremely possible" is just nonsense.
A guy who lives from probabilistic analysis should now better.
I could understand people downvoting my comment (it may be interpreted as harsh even if that wasn't my intention) but... yours? Something weird is at play here.
Silver's point is that their model results in a probability distribution of outcomes and, at the time he was commenting, in 10% of the cases Dems won 19 or fewer seats, resulting in a GOP majority, and in 10% of cases Dems won 51 or more seats, resulting in an overwhelming Dem majority.
That those extremes were approximately equally likely (in CDF terms) tails is willfully missing the point that the vast majority of the 80% of outcomes in the middle between those relatively fat tails involved Democrats winning, which is why the prediction was a 4 in 5 chance of Dems retaining control.
By missing that rather obvious point, D'Souza unsurprisingly comes off as a moron and Taleb also fails to cover himself in glory.
Just based on that snapshot my take is even more grim. Taleb surely knows how probability distributions work. The fact that he sneeringly dismissed basic reasoning like that in the political context he did suggests not cluelessness but malice.
One thing we do know is that Taleb is very smart and very good at making money. And right now there's a lot of money to be spent pandering to the dumb nihilistic and pseudo-intellectual wing of the political right; indeed, the same people that D'Souza has been pandering to for decades.
Unfortunately that is the more common case, and I agree on the author on that that most people take a binary stance on polls.
It should not be the case for regular readers of a website whose main topic is statistical analysis, although I imagine that just before a major election they would see a large number of non-regulars just to check the polls.
If your model can help you predict how it's going to fail, I'm fairly impressed.
I also listen to their politics podcast, and they constantly talk about how to prevent people from rounding up an 80% chance to a 100% chance. One change they made this cycle was to try say 4 out-of 5 chance (which is mathematically the same) to convey both the lack of precision and to make it more intuitive. I think it helps. Personally, when I talked to people about election outcomes, I phrased it as "Based on what I know, I will be surprised-but-not-shocked if x happens, and shocked if y happens" to differentiate between, say, a 20% chance and a 1% chance.
https://fivethirtyeight.com/features/donald-trump-is-winning...
https://fivethirtyeight.com/features/donald-trump-is-the-wor...
> But it’s not how it worked for those skeptical forecasts about Trump’s chance of becoming the Republican nominee. Despite the lack of a model, we put his chances in percentage terms on a number of occasions.
So it would be more accurate to say that he was careful during the second half of the election, after making significant mistakes (both mathematical and ethical) during the first half.
Bur I was just commenting that that the author is spot-on at least in that many readers do not take these explanations into account and they see just 28.6% of winning chance and they automatically convert it to 0% in their mind.
And the explanation is that they don't understand math and probability (or don't want to), whether the author likes it or not.
It makes no sense to me to hold 538 accountable for readers interpretations of the data. If I wander into some realm of science I'm unfamiliar with (say, quantum mechanics) and start making false inferences from it, does that mean the scientists were misleading me?
> The lesson, rather, is that Trump’s campaign will fail by one means or another.
> Donald Trump Is Winning The Polls — And Losing The Nomination
> Our emphatic prediction is simply that Trump will not win the nomination. It’s not even clear that he’s trying to do so.
https://fivethirtyeight.com/features/how-i-acted-like-a-pund...
> saying things with false-certainty and then basking in self-proclaimed genius when luck happened to run in his favor
Something seems to have gone wrong at 538 after their stellar performance in 2012. In 2016 even their Congressional predictions were only 90% accurate. Or perhaps 2016 was simply different from past elections.
Isn't that the very readership they capitalize on?
on edit: grammar was atrocious, still bad now.
I have read read Black Swan and I think Taleb's central thesis is there are valid epistimic gaps in every model, and that suggest there is at least one family of hedge strategies where the goal is to think as creatively as possible about what we don't know, but what could be, so as to reframe that lack of knowledge as a plan to go tromping around in a jungle looking for holes. But to do so requires wide-ranging knowledge across many domains, which is hard to develop.
But overwhelming majority of people don't.
Aleatory is the inherent randomness of a system. You can characterize it but it's often hard to reduce. For example, how much boron impurities are mixed into your steel at the time of manufacture.
Epistemic is things that are knowable but you don't know with much certainty (because it's hard to measure), like the probability that an incident neutron at 1 MeV will inelasatically scatter off a Uranium-238 nucleus and emerge at 200 keV in some direction.
Care to share?
Uncertainty about the outcome of an unbiased die roll is aleatory. You don't know what's going to happen, but that's because of inherent uncertainty, not because you don't know enough about the die. So while you can't be sure what the outcome will be, you can know that each outcome has a 1/6 probability, and you shouldn't expect to adjust that probability as you gain more knowledge.
But uncertainty about the outcome of an election is epistemic, because it depends on a lot of present unknowns. If you run more polls you could reach a different, more accurate probability. Part or all of the uncertainty is because of things you don't know.
I can imagine a world where gambling sharks arrive at the roulette wheel in casinos, wearing hidden google-glass-ish devices, and place bets while the ball is in motion. Unbeknownst to everyone else, their device performed real time computation based on the position and velocity of the ball and wheel, and gave an output of a tight probability distribution (10% chance 25 red, 35% chance 29 black, 40% chance 12 red, 10% chance 8 black, 5% other). The sharks place bets based on these distributions, and make a bunch of money over the course of a few games.
This raises the questions:
1. To the public (not the sharks), all of the possible outcomes of the game have equal probability. Would it be correct to say that all of their uncertainty regarding the outcome of the roulette wheel is aleatory?
2. To the card sharks, they DO have a model with knowledge of game outcomes. So should one say that the distribution output by the device is aleatory uncertainty, and the remaining uncertainty (the fact that there is a distribution instead of a single predicted outcome) in the system is epistimic?
3. #1 and #2 differ in presuming that such a device is present. In absense of such a device, or in general, in absense of a tested model, is it correct to say that all uncertainty is epistimic?
Despite the sniggering at the time, those terms are quite smart.
For different purposes.
One is "probability of rolling a six on a standard die" -- e.g. a systemic property, where we know it's 1 in 6 but we have inherent randomness in how we roll the dice (alea in aleatory comes from the latin for dice btw, as in the famous J.Ceasar quote "alea iacta est" -- well, famous from Asterix at least).
The other is the probability of rolling a six based on what we don't know but in theory could (do we have a standard 6-sided dice? Are we asked to predict an event featuring some bizarro D&D dice we haven't seen? Is it really cubic? Have the edges been treated with a file? How about its balance?)
UQ: Uncertainty Quantification.
(Not, say, Universal Quantifiers, etc).
The crux of the issue taleb has with silver is showing probabilities without making a decision/declaration is cowardly. When asked about making a prediction he says one thing then follows it up with a “ but don’t be surprised if it could be completely random.” And that’s the point - the fact he covers his ass is the problem. Election forecasting is hard so don’t pretend you know something you don’t if you won’t stake something on your predictions. Oh so it could be anything and you don’t want to be held accountable? Then don’t say a damn thing. Basically Silver never wants to be held accountable for his models but always have uncertainty covering his ass as a cop out. so I’m With taleb on this one. Either silver makes a claim and sticks with it or don’t show probabilities and pretend he knows things
And if your data is telling you to be 80% sure, I don't think you should fudge it and claim 100%. Report the uncertainty that you actually have.
I can't think of a similar way to evaluate Silver's election forecasting model. They very clearly aren't independent probabilities, and his model changes significantly from cycle to cycle. Was his model good in 2012 when every state went to where he predicted the likely probability was? Was it bad when his model said Hillary had a 71.4% chance of winning?
You bucket every prediction, look at the outcome, and then confirm whether the favoured outcomes in the 8th decile actually occurred 70% to 80% of the time.
Making a decision/declaration in such a case where the outcome isn't (yet) certain does not mean "being held accountable", it means that you're either stupid, or a liar, or a stupid liar. If you want to call the results of a game before it's certain, then you shouldn't say a damn thing, because anything you say is a lie if you're falsely implying that the result is certain.
If reality is uncertain, then any certain statements/predictions are by definition wrong. Some predictions have more certainty than others, you can stake things on such predictions (sports betting is a great example - if you think that there's a 20% chance of winning but others think that it's 10% or 30%, then there's an opportunity), but it's ridiculous to require certainty where certainty shouldn't be expected.
There are two kinds of things Nate is saying here, and they are related but distinct:
1. Educating readers/listeners about how 20% chances happen all the time 2. Internal analysis and external reporting of how often his 20% predictions were right. If 538 predicts a group of 100 congressmen to all have a 20% chance of being elected and then 21 do get elected, the model did well and Nate's work is worth money to ABC and his reputation is improved (or maintained). If instead 33 of those congressmen get elected, the opposite result happens.
In fact, they regularly do so themselves: https://fivethirtyeight.com/features/how-fivethirtyeights-20...
The quandary that FiveThirtyEight and Silver are in is that the general public is stupendously bad with probabilities. Prior to the the 2016 election, Silver was personally being attacked, in some cases by major media organizations like the Huffington Post, and accused of tipping the scales towards Trump (for some inexplicable reason), since their models were giving him approximately a 1 in 4 chance of winning. Then, in the months following the election, the narrative somehow switched, and the fact that Trump won despite only being given a 1 in 4 chance by FiveThirtyEight meant that Silver was now a hack and the model was wrong.
So now, fast forward to this year's election, Silver is doing his best to drill into people's minds that, yes, their model showed a Democratic win in the house as the most likely outcome, but that absolutely doesn't not mean their model says it is an absolute certainty, or that other outcomes would be unusual. It isn't "covering his ass", its educating the public on how probabilistic statements work.
"Instead, epistemically uncertain events are ignored a priori and then FiveThirtyEight assumes wild fluctuations in a prediction from unforeseen events are a normal part of forecasting. Which should lead us to ask ‘If the model is ignoring some of the most consequential uncertainties, are we really getting a reliable probability?’"
We don't know what we don't know. Any probabilistic model is constructed with a snapshot of the perceived variables. If the model-maker is not able to conceive of variables, then they will not be in the model, which could have asymmetric effects on accuracy of the model.
- It’s not necessary or even useful to define an arbitrary “decision boundary” unless you actually have to make a decision. From a Bayesian perspective, a probability stands for itself: 100% means the event is certain to occur, 50% means you have no idea, and numbers in between convey varying levels of certainty. In reality, 538’s predictions are not true Bayesian probabilities because they don’t take epistemic uncertainty into account, but that’s a totally different issue.
- “Wild fluctuations in a prediction” from new information are absolutely a “normal part of forecasting” - sometimes. If I’m planning to flip two coins, the probability of getting two heads is 25%; but once I flip the first coin, the probability changes to either 50% (if I get heads) or 0% (if I get tails). In the case of Comey reopening the investigation, even if the model had included a probability of that happening, it would have been low and thus wouldn’t affect the overall forecast much. But once that low probability became a certainty, you would expect a sudden swing. The real question is whether 538’s predictions are more swingy than they logically should be (particularly earlier on), but again, that’s a different issue.
To address this, for the 2018 election cycle 538 started displaying their odds as numerical ratios (5 in 9, for example) to try to reduce the aura of certainty.
Did that actually make things clearer to people? Speaking only for myself, I find ratios quite hard to works with and reason about. I certainly couldn't tell you instantly for example if 5/9 is more or less than 4/7 or how big the difference between the two are without first doing some mental arithmetic. Perhaps people in the US are more used to working with ratios since their measurement systems tend to be ratio based.
In 2008 and 2012, the political media was myopically focused on the twists and turns of the horse race and called it 50-50 dead even. Nate Silver's 538 modeled the election as basically stable with only small polling shifts, and the only change in the last month was a steady decline in the remaining time that McCain (or Romney) had to significantly shift the polls.
In 2016, the political media treated the election as if Hillary had a lock on it, and entreated us not to get swept up in the daily swings in the polls. Meanwhile, Nate Silver's model swung wildly with the polls, and Nate went on the media to emphasize the uncertainty of the election. Both candidates were disliked, and a large portion of the electorate was undecided right up until the last week of the election.
I remember the following exchange in one interview:
Interviewer: If you say that Donald Trump has a 33% chance to win, what's the chance that he actually wins? Nate Silver: One in three. I'm predicting a one in three chance that Donald Trump wins in November.
As for the authoritativeness of their presentation, FiveThirtyEight has a really beautiful forecast that presents their 2018 prediction as a probability distribution:
https://projects.fivethirtyeight.com/2018-midterm-election-f...
I don't remember it that way, in fact I remember seeing a lot of very shocked faces in the Democratic camp when the results were beginning to take shape. I don't remember having seen similar confused reactions after any previous US presidential elections, not even after the Bush vs Gore one which was a lot closer in terms of electoral votes.
Back to the article, I think Nate Silver's failure only shows to the general public that electoral predictions are rubbish. Maybe "failure" is a strong word because he genuinely seems to be the best at what he's doing, it's just that the domain in which he's involved is turning out to be bogus. I'm sure that there was a crystal-ball viewer that was the best at what he/she was doing, it's just that crystal-ball viewing turned out to be bogus.
They didn't predict his win, but they did better than anyone else in the mainstream.
Kind of proving my point, as it shows that Nate Silver was the best at crystal-balling the election result. Afaik 30% is still bellow the 50% (or 0.5/1) threshold generally needed to make a decision, as the article also mentions.
In retrospect. In real life, only a few pundits called it from the start (even explaining the mechanics, including Taleb and that Dilbert guy).
Most of the press and pundits was all certain about a specific outcome until the very last minute.
That's what I mean by swingy. Nothing in 2008 or 2012 shifted the polls like that. The polls in those elections stayed within a 2 point band the whole way. That level of volatility was clear in the polls well in advance of the election.
(And, 15% of the electorate was effectively undecided with 2 weeks to go, which is much higher than it was in 2008 or 2012).
Such swings are also what Taleb and the author of this article are criticizing Nate Silver for.
Actually, 100% means the event is "almost certain". That's not an accidental choice of words; it's a technical term which conveys important meaning about how we reason about probability.
https://en.wikipedia.org/wiki/Almost_surely
TL;DR: Infinities being weird as usual.
Imagine throwing a dart at a unit square (i.e. a square with area 1) so that the dart always hits exactly one point of the square, and so that each point in the square is equally likely to be hit.
Now, notice that since the square has area 1, the probability that the dart will hit any particular subregion of the square equals the area of that subregion. For example, the probability that the dart will hit the right half of the square is 0.5, since the right half has area 0.5.
Next, consider the event that "the dart hits a diagonal of the unit square exactly". Since the areas of the diagonals of the square are zero, the probability that the dart lands exactly on a diagonal is zero. So, the dart will almost never land on a diagonal (i.e. it will almost surely not land on a diagonal). Nonetheless the set of points on the diagonals is not empty and a point on a diagonal is no less possible than any other point: the diagonal does contain valid outcomes of the experiment.
Now if someone tells you something is a 90% chance today, and a 10% chance tomorrow, again you don't trust their predictions.
There is a probability associated with "changes in probability". The probability of going from 100 to 0 should intuitively be 0%. The probability of going from 90 to 10 intuitively must be something like 20% (it's like 80 of the 90s "didn't happen", with a probability of 20).
A key point is that you can add this up every day, if something goes from 90->10->90 then it's even more unlikely than going 90->10.
So if you extend that intuition, put some real maths behind it, you can tell whether something is "likely to be a real probability" by the rate at which it changes. And if you have lots and lots of repeated predictions, you can be confident that something isn't a good probability (every timeseries has like a <10% chance of doing exactly what it does, so their combined probability [of being good probabilities] is like 0).
Now Nate Silver accepts this, but says that he's not actually putting a probability on the event in the future, but some non-existent "probability on that event in the future if the future was now". But that doesn't correspond to anything useful or measurable (it's untestable!!), and most people will assume it follows the normal meaning of probability, and it's honestly just silly.
Could you link to where Silver discusses this? I'm interested in seeing his description of exactly what his numbers mean.
What I'm referring to is the "now-cast", but his other two definitions both seem to shy away from saying "this is flat-out the probability we think of the election".
The point is, you can redefine or choose a definition of probability if you want, but if it's less useful than the normal definition (and confusing to people!) then people are free to criticize your work on that basis.
And there's a very useful, testable, mathematical definition of probability that allows us to equally assess everyone's predicting ability, and Nate Silver is dodging it.
If you're interested in this subject, there's a non-mathematical discussion somewhere in Tetlock's book Superforecasting which is interesting in general.
Any discussion of the accuracy of the model (which should be the only thing that admits debate) has to be retrospective. Just saying "well the forecasts changed" doesn't inherently discredit the model, especially because things like elections turn out to be highly sensitive to tiny variations (even something as simple as a rainy election day in a few key precincts can completely flip an outcome).
Which is Taleb's beef with it. The commoner thinks its Silver putting his money on something, but when it doesn't happen he says "well I didn't say _that_".
And it in the day and age of more robust machine learning models, that the 538 models are so volatile is a pretty weak excuse.
Our options aren’t “high confidence converging predictions, or 538”.
The options in the current era for understanding of the electorate’s mood are “overly confident individual polls, poorly analyzed with completely inadequate models”, vs “538-style epistemically humble models accompanied by discussions of their confidence, which can be scored in aggregate after each election”.
I’ll take 538 any day.
538 fails in that most people think that the daily stats are Silver's betting positions. He "predicted the 2008 election."
We understand the difference, but the crying campaign staffers last November did not.
Campaigns aren't relying on Silver’s model.
Also, “politically motivated actors selling narratives that reinforce their preferred outcome largely without data or with cherry-picked data.” Don't forget that option
September 10, 2018 - Hurricane Florence predicted to be category 4 hurricane on landfall, 80% Septemeber 12, 2018 - Hurricane Florence predicted to be category 4 hurricane on landfall, 20% September 14, 2018 - Hurricane Florence makes landfall as Category 1 hurricane
(the above numbers are demonstrative, loosely based on my memory of how events actually happened with Florence)
There is something to be said for predictions about the future based on today's environment, while still allowing for the reality that the environment could change. Predicting a single baseball game right before it happens is just a different kind of prediction than simulating a model that has noisy cross-interacting inputs.
Readers/listeners of 538 need to understand (and Nate spends a lot of time educating about this) exactly what the model is calculating and what it isn't. Nate calls out all the time that the model can only be as good as the polling that provides the inputs. And polls can swing for all sorts of reasons: there's not many of them for a district, only highly biased ones are available, people's actual voting intentions change from week to week.
Am I missing the point of what you're trying to say?
I don’t think their disagreement has much to do with science or mathematics. I think the two guys have big online followings, each leaning heavily toward two different political ideologies, and the two leaders just have to try to beat each other up on a Twitter to decide which side is better.
And he didn't, as far as I recall. I remember him saying something to the effect of "that isn't that impressive, even a simple method looking at polls would have gotten nearly all the states correct". Don't confuse what other people focus on for something he is boasting about.
Isn't that a shared characteristic of any other model dealing with probabilities?
Is the author trying to say Nate Silver gets too much spotlight considering he doesn't guarantee outcomes?
I didn't sell in 2008 and kept buying. I have the same strategy today. Not sure if that also makes me a guru if that's the bar (as the author applies it to Taleb).
Disclaimer: not a mathematician
It is also a cop out to claim "everybody who thinks probabilities based on initial poll responses are not representative of eventual votes" must clearly not understand probability. Maybe I am straw-manning or just don't understand the actual disagreement.
As a poll aggregator? I'm sure it's probably a fabulous tool for past, historical, observed data. The brand itself, probably even more valuable, as a symbol with some sort of social authority in certain circles, I guess.
Anybody who reads my comment history will see I have obviously become a Taleb stan, but his approach seems right to me: "show me the money". Which in his case means mathematical papers of proofs primarily, actual money, second.
I think Silver is actually pretty legit. He was over-hyped, then he was largely panned as a chrlatan but through and through he has been pretty consistent. The Obama elections he rose to prominance but he also did well in the most recent election. The absolute only news outlet saying that Trump had a statistical chance of winning. He hovered around 20-33% at peak while Huffington post told him he was an idiot. Many major outlets didn’t personally attack him obviously but had Trump in low single digits throughout the last weeks
You can read the paper here:
He seems to miss the point of what 538 is doing. Their nowcasts are attempts to say what would happen if the election was today, which throws away the time-dependent uncertainty.
Subjectively, it’s hard for me to believe that an unbiased forecast would truly be so utterly noncommittal until just before Election Day, or indeed that there’s enough data to answer that question, especially seemingly from just one election result. But my subjective impressions, of course, could be utterly wrong. I’m very curious whether or not this is the case.
http://election.princeton.edu/2016/08/21/sharpening-the-fore...
And I think that he was very wrong on that count.
Consider an option on a stock (which is the analogy Taleb is making here). If you buy a 1 year call option on AAPL and tomorrow they announce that they beat earnings by 10%, that's not a huge deal for you. If on the other hand, you owned a 1-week expiration call, it is a big deal for you. That is, your prediction should be less sensitive to changes in environment the further out it is.
I believe he is somehow formalizing this statement, and then showing that Nate Silver's forecasts violate it, but I too don't fully understand his formalism.
This is why Taleb can't make a model that fits Silver's forecasts. Silver could only be confident early on if he knew that people weren't going to change their minds much. But if people don't change their minds much then Silver's forecast shouldn't fluctuate much as time passes. Alternatively, if the forecast fluctuates a lot, it must be because lots of people are changing their minds. But then Silver shouldn't have been so confident to begin with!
But in fact the uncertainty in Silver's model isn't (wholy) caused by the possibility that people change their minds. It's mostly caused by the possibility of polling error. As we approach the election people have less time to change their minds, but the possibility of polling error doesn't change. This why Taleb can't create a model under which Silver's forecasts are rational; he's not taking into account polling error.
If a 538 model says one outcome is 75% probable, it really should come with a giant caveat saying "at current trends assuming nothing out of the ordinary occurs". Taleb's beef is that if that is the case, the 75% does not mean in reality this result is 75% likely to happen.
What is a really interesting takeaway is that if in an event where the outcome you desire based appears unlikely, your job is introduce as much previously undefined or discounted uncertainty. Ideally you engineer a black swan event or at least do what you can to make it happen. This would be fun to model from a game-theory approach instead.
I don't think you can say this. The challenge is that it's impossible to even enumerate all the possible surprises that could drastically swing an election, and incorporating some of them into a model would require pulling numbers out of your ass for how to weight those unpredictable possibilities, and even picking which potential surprises to factor in is a similarly arbitrary decision for which there is insufficient evidence to provide guidance. But none of that means that attempting to factor in such possibilities will drive your model's predictions toward 50%; your predictions could end up almost anywhere in the unit interval depending on the value of the priors you pulled out of your ass, and your final number ends up saying more about your biases than about the state of available predictive evidence.
In chess, if you are behind, John Nunn's two recommended strategies are "grim defence" (if your opponent's advantage is not so large as to make it easy for them to force a win) and "create confusion": create complicated tactical situations and hope your opponent makes a mistake. The farther behind you are, the more appealing "create confusion" gets by comparison.
Silver knows that people's opinions can be changed by events that happen before the election. This is why polls taken early on are less informative about the final result than polls taken closer to the day. Based on historical records, Silver knows how the accuracy of polls depends on the date they are taken, and he weights them accordingly. This process automatically models the possibility that an event could suddenly change people's opinions just before the election. It's taken into account in terms of the accuracy of polls.
[1] https://en.wikipedia.org/wiki/Scoring_rule#Logarithmic_scori...
Another thought: Although the author didn't lie about its meaning, I thought the graph "Stated Probabilities Compared with Average Portions" was visually misleading. At a casual glance, it appeared that 538 was all over the place with its senate predictions. But if you look closer, you can see that many of the red data points are at 0 or 0.5 or 1. This suggests that these data points represent the outcome of just 1 or 2 elections, which isn't really enough to say that 538 is doing particularly badly.
What if Nate Silver were to run around on the football field before each play whispering in the players' ears about the how likely they are to win, according to his latest forecast. Imagine that they believed in him, or at least felt some degree of superstition when he came around.
I just think there's something innately flawed about public forecasting of an election. Mathematically and maybe even ethically. I'm still trying to figure out how to articulate why I feel that way.
It could even be argued that polling itself could influence opinions by introducing bias at a critical moment when someone is being asked to consider who they are voting for. What if that's the moment when they make up their mind?
P.S. https://www.frontiersin.org/articles/10.3389/fphy.2015.00077...
It used to be the case in France that you weren't allowed to publish polls two weeks before an election for just this reason. They eventually scrapped this with the argument that the all 'elites' had access to internal and private polls anyway so all you where doing was denying the 'people' the same access to information that the 'elites' had.
It could even be argued that polling itself could influence opinions by introducing bias at a critical moment when someone is being asked to consider who they are voting for.
This is a pretty well known effect, to the point that many political strategists try to avoid publishing or drawing attention to polls showing their candidate too far ahead in case it makes voters complacent and decide to stay home. The ideal situation is to make people convinced the polls show the candidates tied 50-50 going in to the election and that 'your' vote could easily swing the whole election.
However I totally dig what 538 does and appreciate their analysis. I understand that there is a complex underlying model with many parameters taken subjectively. The results of Monte Carlo runs of this model are very interesting to me. The fact that they describe distribution of outcomes and not just a single number is important property of a Monte Carlo model and I would have an issue if that wouldn't provide that info. If I would like to analyze complex event that has lots of moving parts and uncertainties I would also build a model to see what kind of distribution I would get. The fact that somebody published results of their model is quite useful.
The analogy would be European and American weather models - no one say that their results are exact and everybody understand that uncertainty in result comes from inherent uncertainty of initial state as well as shortcuts and approximations each of the model takes. No one says that the results of these models are useless because of that. And everybody finds it valuable if weather forecast predicts rain tomorrow with 40% chance. So I don't see how political prediction is only valuable (by the words of the author) if it comes without probabilities attached to it.
Yes, this is good. It means that you can evaluate their accuracy to a greater degree than % of times correct.
> Further complicating the issue, these predictions are reported as point estimates
They've learned from this, and now show their distribution of results very clearly. However, there's nothing mathematically fraught about reporting a prediction as a point estimate.
> The problem is that models are not perfect replicas of the real world and are, as a matter of fact, always wrong in some way.
...yes? And so what? This is a fully general counterargument to all of science.
> Predictions have two types of uncertainty; aleatory and epistemic.
Don't say things like this with 100% certainty when your source itself says "the validity of this categorization is open to debate", and also when your source is Wikipedia.
> However, as you can see, there is still a noticable variation of 2–5% of actual proportion to predictions. This is a signal of un-addressed epistemic uncertainty. It also means you cannot take one of these forecast probabilities at face value.
Models aren't perfect. Also, 538 is very careful to address their epistemic uncertainty! They also try very hard to not change their models significantly after they publish them, in order to avoid letting their personal biases tinker with the results.
This post goes out of it's way a to avoid mentioning that there are pretty well established ways of measuring the accuracy of predictions, instead using things like "look, it's not quite a line" and "look, the line went up and down before the election".
If you're interested in actually reading about this from a mathematical viewpoint, read Taleb's paper, which has some actually interesting thoughts on why 538's algorithms are too eager, or read Madeka's paper, which includes comparisons of 538 to other people trying to make predictions, and finds that actually, they were better than most other news sources.
The way he talked to renowned classicist (and feminist, which seemed to be a problem for Taleb) Dame Mary Beard in another spat was disgusting. Trying to impose statistical certainty on ancient history ended up making him look pretty dumb to everyone but him.
Btw, I read 3 of Taleb's books, 1 from Nate Silver. Both are pretty good. Taleb is the more original, but also more often-objectionable.
So what should the bet be? As far as I can tell they're not disagreeing on anything that can easily be reduced to a bet
Do this for each election.
That being said it would be a fun little exercise to see who much money would be made if they bet had $1 on every race 538 predicted in 2016 and 2018 using the model odds.
Note: I'm don't think Taleb has a wrong idea. It's just that he attempt to fight with Silver over it is based on confusion. Taleb has written important and interesting books and papers. He has always been very opinionated and aggressive. His feuds don't produce debates that are worth following.
Election Predictions as Martingales: An Arbitrage Approach, Nassim Nicholas Taleb Quantitative Finance. 452 (1): 1–5. doi:10.1080/14697688.2017.1395230 https://arxiv.org/abs/1703.06351
https://www.businessinsider.com/taleb-every-single-human-bei...
>>> "every single human being" should bet U.S. Treasury bonds will decline.
>>> Short the S&P vs Long Gold, in a 5 to 1 ratio. By gold Taleb means a basket of precious metals including gold.
The second trade would have lost $125k for every $10k invested in gold and $50k invested in SPY.
Did he miscalculate the risk of a black swan event...?
Nate Silver took political polls and hunches, which were far less accurate, and came up with something significantly better. Even in 2016, Fivethirteight's model showed the odds of Trump winning were not that remote versus major media outlets "guessing" the race was all but over.
Election prediction models will always be inherently unstable in my mind. It's pretty much impossible to predict the when, if, how, who of an Anthony Weiner type scandal.
What do you mean by a business cycle? If his fund just goes until it makes it big and then he stops, is that not a business cycle? Did he continue? I don't know. It just seems weird to ask about a business cycle in the case of a strategy that is very long played, waiting for unpredictable results. The only thing that matters is can you play long enough to hit that result if that is your game? It seems clear that it did win, or else he'd probably be focusing more on it until he did? It's not like a regular business with actual output or production that can be easily replicated and measured, it's just a play till you win or lose deal and I'm sure there are others that did similar things and lost, just like many more companies are dead than alive and we only see the alive ones.
I understand wanting to see the numbers, but this is just one sample for a strategy.
> Hyperinflation bet that could very well not work but if it does "you will never fly in a public jet again."
So it's a bet on a small probability with very high payout. This approach worked for him in the past.
Silver's predictions fluctuate a lot as time passes. Taleb is assuming that all of these fluctuations are caused by the electorate changing who they're going to vote for. Then he says "if you knew people's opinions were as volatile as this, you shouldn't have been so confident earlier!", or words to that effect. If lots of people will change their minds, then early polling won't be very informative about the final result.
But the fluctuations in Silver's predictions are actually caused by his changing certainty about how accurate the polls are. Silver does take into account that people might change their minds, but this isn't the cause of most changes in his predictions. So people actually change their minds a lot less often than Taleb's model of Silver's model suggests. This is why Silver can make strong predictions early on.
[1]: http://www.fooledbyrandomness.com/pp2
[2]: https://medium.com/opacity/the-syrian-war-condensed-a-more-r...
[3]: https://medium.com/@ameraidi/nassim-taleb-is-wrong-on-syria-...
It doesn't have to be a binary choice. Knowing the distinction between 65% and 80% improves your betmaking capabilities by allowing you to hedge, in some cases guaranteeing a profit [1] regardless of the odds and how they change.
[1] https://help.smarkets.com/hc/en-gb/articles/115001431011-How...
The biggest and most unpredictable factor in US (and British) elections is their winner-takes-all system, with the second layer of the Electoral College in the US which does not account for popular votes either.
I listen to the 538 podcast every episode, and I think I agree with cm2012 that this is a case of mixing up the "nowcast" with the actual forecast.
True, the epistemic uncertainty is limited -- it's based on past elections, because hey, 538 works with actual data. One could argue that rather than saying "71.4% Clinton 28.6% Trump", a better model would have said "71.3% Clinton 28.4% Trump 0.2% Election is cancelled due to nuclear war / natural disaster / etc" but I don't think anyone sensible is interpreting the model as excluding such extreme outcomes.
My favorite Nate Bronze takedown remains Carl Diggler[1].
[1] https://www.washingtonpost.com/amphtml/posteverything/wp/201...
1. It conflates the forecast with the now-cast. The former is a prediction of what will happen on Election Day (whether that day is six months away or one day away). The latter is a prediction of what would happen if the election were held that day. In theory, those two would converge to the same value on the day of the election, though in practice, that wouldn't actually happen due to computational differences.
2. Silver never says that he "should be judged" by [just] the final result. In fact, he's gone out of his way to say that. The problem is, that's the only point on the graph where we can compare a model (predicted value) to the actual (observed value). There isn't an election on any of the other days, so even if both agreed to look at the now-cast and use that as grounds to evaluate the model, we still would only have one datapoint which is nonzero in both dimensions. In other words, we have a blue line, yes, but we only have a red dot. Taleb wants to extrapolate the red dot into a horizontal blue line, and then use that to judge Silver's model, which is ridiculous.
If I then make another prediction that's always 1 year out, you can decide how good that prediction is and start to make a curve of how good my predictions are by how far out they are.
Taleb is saying that the blue line is so incorrect it's meaningless and deceptive and Nate should product a red line which is a real probability that he can actually be judged on.
The quandary that FiveThirtyEight and Silver are in is that the general public is stupendously bad with probabilities. Prior to the the 2016 election, Silver was personally being attacked, in some cases by major media organizations like the Huffington Post, and accused of tipping the scales towards Trump (for some inexplicable reason), since their models were giving him approximately a 1 in 4 chance of winning, since everyone knew it was a certainty Clinton would win. Then, in the months following the election, the narrative somehow switched, and the fact that Trump won despite only being given a 1 in 4 chance by FiveThirtyEight meant that Silver was now a hack and the model was wrong. This same criticism has been leveled by many right-wing media personalities, who had been trotting out Silver's 2016 predictions leading up to the election, as well as by the New York Times, whose own models were giving Clinton nearly a 90% chance of victory.
So now, fast forward to this year's election, Silver is doing his best to drill into people's minds that, yes, their model showed a Democratic win in the house as the most likely outcome, but that absolutely doesn't not mean their model says it is an absolute certainty, or that other outcomes would be unusual. It isn't covering his ass, its educating the public on how probabilistic statements work.
This is something Taleb should know. But Taleb has proven himself to be a hack and an intellectual parasite. He hasn't actually produced any ideas worth discussing since "Fooled By Randomness" which was nearly 20 years ago. Basically all of his works since then have been rehashing the same idea, or, as with Skin in the Game, pseudo-intellectual ramblings containing little to no interesting ideas, and what ideas do exist have been rehashed a million times by others.
Since then, he has taken to making ad-hominem attacks and nonsensical critiques on others in the public view, often on Twitter. I hate bringing up Trump, as it happens far too often in discussions these days, but some of his methods of "debate" are strikingly similar: strange monikers ("klueless Nate"), schizophrenic jumps between topics/arguments, often indecipherable language, and a ridiculous volume of output. All of this combines to mask the fact that most of the time, he isn't actually saying anything, or at least anything sensible, but by shouting loud enough, long enough, and purposefully making his point difficult to identify, it almost becomes impossible to argue with him, and a non-insignificant number of people will assume he is right.
Interestingly, I've also found many of his supporters to be similar to some of Trump's or other figures such as Musk. There is a cult of personality built up there where they have established the person as a visionary first, and therefore anything and everything they do or say is correct. Many of them parrot back the same, shallow catchphrases of his, that rarely contain anything insightful but are general enough that they can be thrown out in any situation: critics or dissidents are "Intellectual yet Idiots" or don't have "Skin in the Game". Note that the arguments of these critics or dissidents are almost never addressed.
Unfortunately, the spat between Taleb and Silver (which actually started back prior to the 2016 elections), is exhibit A in why Taleb is an intellectual charlatan, and shouldn't be given the time of day. Silver isn't some perfect person either, but in general he has been a reliable voice in the realms of political prediction, fairly forthcoming on shortfalls, and his most damning criticisms all seem to be built upon critical failures in the understanding of probability, or a willful misrepresentation of his results.
Because the polls were wrong. Or, at least, incomplete. It doesn't matter how good your algorithms are if the data isn't there. And 2016 proved that 538, and really most media outlets, don't have a grip on America.
In other words, the existence of a black swan event in that particular system didn't really matter.
Silver is a glorified modern fortune teller, who seems to have deluded himself into thinking that simply having data means you can prescribe meaningful probabilities to the future. Assigning probabilities to a coin flip is easy. Figuring out the most likely winner of a game of baseball is a little harder; there's a whole movie about a guy who was basically doing that in 2002. Politics is a completely different field. You can't possibly assemble all of the data necessary to be remotely confident in the probability of outcomes. And let's not forget unpredictable black swan events.
But the readers love it. And, especially during the elections, the media would bring him out like a golden boy computer whiz, because he was saying the same things they were, but he's really smart and has the data and magic algorithms to back him up. And he was right that one time 8 years ago. He's about as pointless as Sean Hannity, with the one exception that, at least recently, Hannity was actually more correct about the future than Nate. It doesn't matter if you are right or wrong, if the reasons are wrong.
FiveThirtyEight Prediction: 48.5% to 44.9% Real outcome: 48.2% to 46.1%
I agree that the horserace coverage is silly (especially given how polling operations use sampling and such), but its better to look at aggregate trends in polling than to ridiculously overcover outlier polls.
I also certainly don't think its fair to say "28.6% chance he wins" is a "massive win" prediction. Silver's take was regarded as indefensibly right-leaning and attacked.
It was predicting a large electoral win as the median scenario, but only a 70% chance of Clinton winning at all, because discrepancy between actual outcomes and polling were likely to be correlated across states, and many states were very close. This is where Silver differed from the other forecasts that were saying Clinton 90%+ based on the (historically false) idea that differences between votes and polls were independent across states.
> You can't possibly assemble all of the data necessary to be remotely confident in the probability of outcomes.
You often can be more than remotely confident, but it's true that the national result of the 2016 Presidential election wasn't one of those times. OTOH, Silver’s forecast reflected that, since ~70% is extremely low confidence.
Nate was literally one of the core people working on Sabermetrics around the time that Moneyball was written.