Show HN: Applying Machine Learning to March Madness
github.com
github.com
But if you really want to try to solve this problem, you're going to need more granular data - play-by-play, per-play lineup, injuries, player tracking if you can get your hands on it. The stats you are using are very lossy summaries of the season, so they aren't very strong predictive features.
From a technical POV, consider bucketing some of those per-game stats (e.g. 4 binary features representing which quartile the team's stat falls in compared to every other team that season). This can help to adjust for year-to-year differences. Work with pace-adjusted stats if you have access to them. Find a baseline accuracy by picking the simplest possible non-ML strategy and measuring how accurate that is (e.g. what is the accuracy of a model that always picks the better seed or W/L record?).
You need to adjust your training data to only use data that would have been available at the time of the game - if W/L or PPG includes data from future games, this is a form of data snooping and will probably give you results on your test set that won't generalize to the real world. Time-series snooping is a very easy mistake to make, but it's crucial to avoid it in order to build a good model.
Interesting work, thanks for sharing!
The odds coming out of Vegas are usually priced correctly. Sports markets are very efficient—although perhaps not as ruthlessy efficient as public equity markets. I would imagine there are still syndicates out there that are the “RenTec of sports betting” and just printing alpha.
This is wrong. Many of the games are not even close to 50-50.
To put the odds of picking a perfect grid by hand in rough perspective, it would be like winning the lottery, boarding a plane that crashes into the ocean, being the only survivor floating in the ocean, but then are struck by lightning 3 times and yet you live, only to then be eaten by a shark, all in a single day.
That makes me wonder what odds you'd get at the bookmakers for a perfect winning grid.
Edit: it was a billion.
https://genius.com/Warren-buffett-billion-dollar-march-madne...
XD
https://adeshpande3.github.io/adeshpande3.github.io/Applying...
http://www.sloansportsconference.com/wp-content/uploads/2018...
There was a mainstream article written about the tech a while back, which had GIFs demonstrating simulations, but I failed to find it. Essentially, it would model what the players did versus what they should have done, optimally.