A mathematical formula that predicts US elections with 87.5% accuracy
turingbotsoftware.com
turingbotsoftware.com
One can fit an elephant with four free parameters. That's the story of these curve fitting exercises.
I feel like the message I'm supposed to take from this demo is "Look how good TuringBot is! They can automatically find a function to match this data!" But the actual message I'm getting is "Symbolic regression is too hard for TuringBot!"
Since you have a literal dependency between one data point and the next, you can't train your model using randomized data. So the default that people jump to is to segment their data with earlier data being the training set, and later data being the test set.
If your data is the result of a very well understood and controlled process, this works fabulously (see ARIMA and variants). But the more that the data relies on significant noise, whether that is pure randomness or Brownian motion (very likely the case with elections), that methodology breaks down spectacularly. All you end up doing is overfitting to your test set instead of your training set.
[0](https://en.wikipedia.org/wiki/The_Keys_to_the_White_House)