Assigning a probability to a one-off event is enumerating all the ways it could happen, all the ways it could not happen and assigning a probability to each of those ways. Obviously there is a lot of guesswork; but if you need to make a decision based on the future that approach gives you a much better chance of making a good decision. In practice an event will be made up of components that are more predictable than the whole, and some real uncertainties. There is a lot to be gained by thinking hard about the situation, and assigning a probability will do that.
Silly example - how do I estimate my risk of falling climbing up a set of stairs that I've never climbed before?
* Baseline risk of tripping - I have a lifetime of data.
* Increase of risk being on a staircase - I have an area I want to put my foot on (a specific step) that is about 1/3 of the area that my foot usually falls in, so that increases the risk by an amount that can be reasonably estimated.
* I will watch my foot - maybe an order of magnitude improvement in precision.
This is enough to let me estimate the risk of carrying a bulky object (that obscures my view of my feet) up a staircase. I've isolated the uncertainty (how much does visual observation change the odds) from the certainties (areas, background rate).
Now I can take that to several experts who will identify new mechanisms and tighten up my estimations on how big a deal the components are. In this way - even though the final % I come to would still be a bit arbitrary - it is starting to become a summary of what a large number of people think about the inputs to the problem and their relative magnitudes. Being able to communicate all that thinking with a single number is a miracle in its own way.
The data might favour one candidate, but even assuming it is unbiased and representative, it is only a random sample, and there is a chance it could be randomly wrong.
Kidding aside, your observation is correct, you need to perform repeated predictions. 538 does just that. They keep predicting a whole lot of outcomes, so you can check their track record. They even have a challenge for the audience, where you can record your own predictions and compare them with their corresponding predictions (see for example [1] for NFL games). The scoring of the predictions uses the Brier score [2], which is just a version of R-squared for classification problems.
In the case of the November 2016 election, other prediction sites were giving the Trump team a 1-2% chance, 538 was giving them a 20% change. Considering that, 538 comes out as the clear winner.
Which brings us to what I consider to be the best measure of prediction accuracy (my own invention, it doesn't have a scientific name yet): you can compare two sequences of predictions (let's say yours agains 538) by simulating bets at some mid-odds level. For example if I say Trump's probability of winning 2020 is 45% and you say it's 35%, then I'm willing to give you 2-to-3 odds, while you're willing to take that. We then see who stays ahead most of the time in this betting simulation.
For what is worth, here's a recent survey of different scoring types for predictions: "Assessing the performance of prediction models: a framework for traditional and novel measures" [3]
[1] https://projects.fivethirtyeight.com/nfl-predictions-game/ [2] https://en.wikipedia.org/wiki/Brier_score [3] https://www.ncbi.nlm.nih.gov/pubmed/20010215
One common mistake I've seen is people mistaking that 20% as the expected vote ratio. It definitely wasn't; that was hidden behind an apparently too-well-hidden toggle (buttons above the graph), and was floating around 45% Trump and 48% Clinton, with overlapping error ranges.
But the betting odds depend on how much knowledge you have - for someone with perfect knowledge the odds may be 0 or 1, if you know nothing you might guess 50:50 for practical purposes.
So, for example, when 538 said that Trump had a 30% chance of winning, its hard to say how accurate that was given it was a one time event. But if you look over all of their election predictions, and look at those they predicted to win 30% of the time, about 30% of them should have been victories (with some margin for error, of course). If 50% were victories, or only 10%, then maybe that would indicate their models aren't doing so hot.
538 also predicted lower chances of victory in the weeks and months leading up to that, then adjusted it on election day vanishing away their earlier guesses. This isn't wrong on the face of it, but people infer from the final datapoint to the conclusion later that same day about 538's accuracy.
What I am saying is: 538 needs to keep a historical record of their predictions.
It has not changed. They did this the whole way back when it was an independent blog before the NY Times (later ESPN then ABC News) picked them up.
They are literally charted on the page with the current prediction, e.g., for the 2016 Presidential race (which is locked on the last forecast):
https://projects.fivethirtyeight.com/2016-election-forecast/ , under “How the forecast has changed”, it charts the forecast (by default, the win probability) for the entire time the race was being forecast.
[1] https://projects.fivethirtyeight.com/2016-election-forecast/...
It's worth noting that you are completely wrong: not only do they keep such a record, they also provide a handy chart of it on the latest prediction page so you can easily see exactly how the prediction changed over time.
Edit: For example, here are the headers on the file named presidential_elections.csv
year,office,state,district,election_date,forecast_date,forecast_type,party,candidate,projected_voteshare,actual_voteshare,probwin,probwin_outcome
These are the actual forecasts they made for presidential elections going all the way back to 2008.
- Web: https://projects.fivethirtyeight.com/2018-midterm-election-f...
- JSON: https://projects.fivethirtyeight.com/2018-midterm-election-f...
You are correct they do not publish the exact math of the forecasts, but you can see how the forecast changed over time.