Interestingly, a frequentist approach (i.e. just looking at what happened in the past, producing a rough regression, and then assembling ensembles of such models), is likely to get you +EV in markets that don't get frequented by big hitters. It'll help you in your high school football league, for example, but it won't help you make money in the NFL markets.
Benter has written about his approaches, and has hinted at various ways of working he used.
The frequentist approach he took most likely still stands up in Hong Kong because there are only two tracks, variance is limited, the horse population is limited etc.
That doesn't apply anywhere else in the World, or for almost any other sport.
There is a whole World out there that you are yet to discover, but it's unlikely the syndicates will welcome you. More money in the market on the other side of the table is always welcome (that's liquidity, always helpful for the market to be bigger), but more money on your/their side makes the job harder.
Always remember to only bet what you can afford to lose, and good luck!
The bayesian approach says the data is fixed, and the probabilities might change.
If I look at the stats for the NE Patriots, for example (http://www.nfl.com/teams/newenglandpatriots/statistics?team=...) I see that their first down conversion rate is 30/56 so 53%.
Imagine I am watching a game and I am betting in play. I am offered odds of 1.9 (decimal odds) that they will convert the third down they are about to try and convert into a first down.
The frequentist approach says given the implied odds of past behaviour is 1.88 (we can convert percentages into decimal odds by dividing 100 by the percentage, so 100/53 = 1.88), and I am being offered 1.9, I should bet! Kelly says I should bet 0.78% of my bankroll, as I have an edge here.
Now, does that make sense to you? 30/56 is what happened in the past, and we're using that as an indicator as to what to do next. Would you take that bet?
The problem with this approach, I think, is that frequentist approaches whilst practical assume there is an underlying probability we can uncover by measuring it.
The Bayesian approach (in simple terms), says we can't be that precise, and the probabilities change over time based on the context. This makes more intuitive sense: the probabilities in poker are fixed and calculable, it seems to me they are much less so in NFL games.
In the Bayesian approach, we broadly need to think of a probability distribution and understand our confidence interval, and we use priors and observations to help us calculate both.
Doing some maths we might say the chance of the Patriots getting the third down conversion is with 95% confidence the chance of between 51.5% and 54.5%.
Well, now the 1.9 on offer isn't quite so sweet - it's within the confidence interval, albeit off to the edge.
Getting to that distribution and narrowing your confidence interval (it would be great if we could say it was 52.8% to 52.9%, for example), and then figuring out how to use Kelly accordingly, is relatively state of the art.
Doing this in the NFL might be tricky because the data sizes are relatively small - the confidence intervals might be too broad. Also, the frequentist approach is provenly useful in some situations: Bill Benter is richer than either of us, and I don't believe he ever used bayesian statistics.
People often think of gamblers as slightly grimy/shady characters with a gold chain and a wad of bills in their hand. That might happen, but all the ones I speak to spend their weekends reading PhD theses from maths and finance departments where people have been trying to figure out this stuff. I hope this answer gives you a flavour.
Other jurisdictions not so much. There are various scrapers available to obtain generic race card information[2] and APIs from some bookmakers/exchanges.[3]
A couple of years ago I read an article lamenting the lack of open data in horse racing, which still largely holds true today. Unfortunately, I can't find the article now (it was either on the Paulick Report or Bloodhorse website), but here's a more recent one from May this year calling for more open data.[4]
In fact, a question was asked on HN a couple of years ago about an open horse racing database and as far as I know the answers there are still valid.[5]
In general, there is data available for US, UK, IRE, Australian racing (the main jurisdictions) but they have to be paid for.[6][7][8][9]
[1] https://racing.hkjc.com/racing/english/index.aspx
[2] https://github.com/4A47/rpscrape
[3] https://docs.developer.betfair.com/display/1smk3cen4v3lu3yom...
[4] https://www.bloodhorse.com/horse-racing/articles/232487/thor...
[5] https://news.ycombinator.com/item?id=7779117
[6] https://www.betwise.co.uk/smartform
[7] https://www.proformracing.com/
Several orgs, including Equibase (US-based, the gate keeper of a good portion of handicapping data) will regularly send cease and desist orders to people who attempt to automate aggregation of data even with free, publicly available content. That's at least half the reason PDFs are used when customers purchase data access, to make aggregation harder (you should see some of the white space, character encoding fuckery they use to throw off aggregators).
I suppose some of this often depends the quality of the data as well. Most data entry happens at the track during the race by a human, none of the data collection about races or the horse stats are collected by a computer, it's 95% hand entered. That also goes for pedigree information and other statistics including medications, weights, etc. And 100% of that is usually self-reported.
Much of the current handicapping in the industry is everyone trying to protect their personal mountains of data. Tech-minded people would love to provide open, controlled, API services so that people can do what they will with our mountains of data. But "giving it away for free" is a non-starter for the good ole boys at the top..
I was involved for a number of years with a UK based horse racing ratings service (handicapping if in the US). This service used to license their base data from the Press Association[1] and then run algorithms on top to produce the ratings.
There's certain things I can't say due to NDAs which are probably still in effect, but the cost of licensing this basic data was in excess of £10k per annum. So, unless you were a serious bettor or were looking to operate a service of some kind, it's beyond the pocket of most individuals.
Timeform in the UK also license some of their own proprietory data, via an API[2]. They've published some pricing on their website and you're looking at between £6k - £12k per year. This is just to access data which is available via their website for a subscription fee of £75 per month, but via their API.
There's even a specific UK organisation which apparently has the permission from the British Horse Racing Authority to officially licence key racing data. This is who sells the data to bookmakers, form guides, racing newspapers etc. They have a rate card published on their website.[3] Private, pro-punter? £8.5k per year please.
It's a bit of a rort really. Most of the data is "freely" available online or in the racing press, but if you want to access it any useable format, either build a scraper (good luck with staying on top of the website changes) or pay a stack to access things programmatically.
[1] https://pa.media/racing-betting/horseracing/
Almost all tracks publish result charts online for free along with race videos. If you want free, why not compile the data yourself? How long would DRF or Equibase exist if people could access their data for free?
Also, it's important to make the distinction between editorial content (analysis, predictions, subjective descriptions of a horse or jockey performance) and empirical information (horse weights, medication, surface conditions, weather, placements, jockey-horse combo win-rates, etc).
The DRF sells its speed ratings as well as analysis of pedigree and past performances. There's value in that and it definitely justifies the cost of their publication and the other publications that perform similar work.
The critical issue with your stance is that users have no options to aggregate their own data easily. The free PPs Equibase offers have been scrapped before and I know of several specific instances where the creators of those scrappers were sent cease and desist for collecting the information Equibase otherwise provides for free. Even to Github to remove the repository that contains the code.
I'm not advocating scrapping (please don't scrape sites like that) but there isn't any industry interest in providing modern consumable data. Wouldn't it be in Equibases best interest to put that information behind an API and sell access to the public? The industry actively discourages using publicly available data.
I really do not care about the likes of DRF or Equibase and how long they will or won't exist. I think it is upon the industry itself to ensure this data is available free and easily accessible. Look at Hong Kong as the alpha example. Loads of free data, huge betting turnover, well funded industry.
DRF makes racing data easily accessible. If it was left to the tracks, which are independent entities (unlike NFL/NBA/MLB), an horseplayer would have to compile past performances from dozens of sources. The fields of a single day's race card may have run at 30 or more individual venues, in aggregate. Even if that data were free (well, the result charts and replay videos are already free, so technically this is already possible) if would take a ton of work to assemble it all in a digestible format -- which the DRF does for 6 bucks.
I don't believe HK offers free data that is not available from American tracks. There is no API, the result charts are less detailed than American tracks. If info was so freely available to everyone, how would someone like Bill Benter gain such a huge advantage? Why wouldn't he replicate his methods in the US? Probably because the US makes MORE data available.
Benter used this himself, as detailed in his paper 'Computer Based Horse Race Handicapping and Wagering Systems: A Report' - https://www.gwern.net/docs/statistics/decision/1994-benter.p...
I'm surprised he is openly talking about everything, but I suppose many others have caught up and there's not much to lose ...
Then he'd go to Belmont in the Spring and basically live at Saratoga in the Summer. He'd watch workouts, assess the horses, watch the trainers and jockeys and gather various datapoints. I was skeptical at first, but he at a minimum made enough to live in Saratoga for the meet, which is not cheap, and funded alot of his toy purchases from his winnings. He'd call out of the blue give me picks that were good winners about 50% of the time.
Basically, there were two ways he made money -- he followed his dozen horses for the high quality stakes races, and was able to eliminate shitty horses for the lower quality races based on trainer or workout performance. For the low quality races, he would hit a few wins/exactas a week (exacta = bet on 1st and 2nd place), and for the higher quality races he would do more exotic bets (Pick 6, trifecta, exacta)
This guy loved horses, the track, and the people around it. He was a widower and had lots of time on his hands. Definitely a labor of love that kept him sharp for a long time.
The big syndicates also aren't that secretive anymore. The founders and employees have got too rich (some are billionaires), and they have had to build a profile to hire (if you attended an elite uni in the UK, they recruit there...the largest syndicate in the UK has hundreds of programmers now).
And Bill has been talking about this for years. I don't understand why but he was publishing papers and giving talks in the 1990s. The impression I get is that he wants other people to know he is smart, and already has more money than he needs (this was true even in the 1990s).
Something that most people here won't get though: the models don't matter. The largest syndicates hire clever people but the real money is made in finding liquidity and people to take the other side. The big UK syndicates got large because they worked with bookies in Asia (the largest syndicate is run by someone who ran an Asian book) who needed people to take risk.
Btw, these websites already exist but no-one on there makes any money. You don't need a lot of capital to make money in sports betting (because you turn your capital over so many times a year), there is no reason for a good sports bettor to use those sites to publicise themselves.
[1] http://www.puntingform.com.au/account/systems-dashboard/#top...
https://www.espn.com/blog/playbook/dollars/post/_/id/2935/me...