How I'm Predicting Baseball Outcomes
blog.zachschnell.com
blog.zachschnell.com
I honestly don't. (You are very, very far off the correct methodology.)
But maybe this is a good place, I guess.
http://www.oakton.edu/user/4/pboisver/AABaseballMath.html
I am not trying to be a jerk. But ask yourself: If it were that trivial, why wouldn't it be a snap to crush online sportsbetting?
Just a random article. I've met him through poker. He had a radio show for a while, and he stated that if you have a 2% edge over a bookie, you can easily make a million dollars a year. Anyone who thinks they have an edge and isn't also very wealthy is simply mistaken.
Baseball is basicly a lineup versus lineup sport, right? Or in terms of the link: Q per team is variable between matches. An additional modelling level should estimate Q per team per match, before estimating Qx - Qy.
Without Cabrera, DET has a harder time winning. Same for the big bearded guy.
As a sidenote: some guy posts his code (+) and gets upvoted (+). He's not aiming to change the world, or to become a 100-million firm. He wants feedback and to improve his approach. We can give him that without the "aararrrgh, but Terrence Tao is much smarter than you"? Otherwise, just don't upvote :)
You might also be interested in checking out the book The Signal and The Noise by Nate Silver (He runs the FiveThirtyEight political/data blog that notably predicted last year's election results with great accuracy.)
I actually just started that book and so far so excellent.
PECOTA is alright, but not really his anymore. cwyers took it over at BPro. It's also tough to say if it's one of the best for any number of statistically boring reasons (read Phil Birnbaum's blog for more info).
Birnbaum is one of the best bloggers on advanced stuff. Tango's blog is the best period as anyone good participating in the "open source" movement so to speak is there pr shows up when they write good stuff. Until you're aware of the work to date, you will be spinning in circles with awful biased errors.
i.e. the author regrouped the data to come up with the same result
To start with, you can't 'predict' result of a single game. You can have some advantage (or more likely, disadvantage) over the quality of the betting like a bookie gives you.
That said, there are a gazillion other data points that are more granular and would provide more predictive value than simply the binary result of prior games (especially if those games took place 12+ months ago).
But at its least this was good coding practice.
I'm not a statistician, so I can't really speak to how 'valid' the analysis is, but I'd be curious to see how it does in different tests--unless I'm misinterpreting, the biggest check so far was done on 2012, which is exactly what was used to train it. It would be interesting to see what happens if you train with half of 2012 and test the second half. Or check 2011 (do you predict an end-of-the-year collapse, allowing my Cardinals to sneak in again? ;) ).
I pitched 5 mph less due to mental effort? Well, now you don't get to throw again for 5 days. Or you walked 15 batters in a row to try and get taken out of the game. It's not going to make your night easier, you sit there till the game is finally over. No clock to run out. Just outs.
Most of all: baseballers are fairly superstitious. They believe in streaks, lucky rituals, jinxes, and the gamblet's fallacy (being 'due' for a win or a loss). So some serial correlation could be a self-fulfilling prophesy.
Still, I'd expect tons of other available team stats to outperform the last N game results in predictive power.
Baseball stats pre-molded into a nicely workable form, available from your handy R interpreter.
Baseball fan in me says that is correct but statistician in me would like to see more models like yours to quantify it.
jerf@jerfhom:~$ python
Python 2.7.3 (default, Sep 26 2012, 21:51:14)
[GCC 4.7.2] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> import random
>>> 94.0/(94+68)
0.5802469135802469
>>> winp = 94.0/(94+68)
>>> games = []
>>> for x in range(50):
... games.append('w' if random.random() < winp else 'L')
...
>>> ''.join(games)
'wLwLwwwwwLwLwLwwwLLLwLwwLwLLLwwLwwwwwLwwwLwwLLwLLw'
In my full simulation of 162 games, the longest streak was a 7 game losing streak, despite the higher win percentage. Of course you'll get different results each run; my next run produced a 9 game winning streak, which some quick Googling suggests is in line with what happened in 2010.Combine this with the fact that real play is not drawn uniformly (you may play a much worse team against which you have a much better win percentage for several games in a row) and I don't see much need for some sort of meaningful, statistically-predictive "streak" to explain game results.
And I understand exactly where you're coming from. This is very preliminary, and if anything it was good coding practice for me. Though I very much intend to incorporate more significant factors like the lineup, the opposing team, and their history.
Based on these models, you should have some good examples of selection bias, and see how the model changes based on what you are not testing for, but what is implicit in the data (since data is merely a set of samples of data generated by one iteration of the (unknowable to some degree) true talent functions for each team (player, lineup decision, injury, close call by an ump, etc.)
If you're interested in going down the rabbit hole, there's tons of people who can show the way (and they're nice! At least tangotiger is way nicer than he should be in listening to people who have put no effort in understanding what is good and what is beginner's blind bliss)
Hot and cold streaks are just random variance, so is whether balls are hit within reach of fielders or safely out of reach, given a certain contact quality (ground ball, fly ball, infield pop up, or line drive all have vastly different tendencies to fall for a hit - line drives ~.600-700 babip if I recall, FB ~ low .200ish, GB ~ .300, pop up 0ish?) point is these are all known, to se degree, given the historical data.
If anyone wants to explore this stuff further let me know & I can point you to the right spots to help a specific interest?
I would usually presume against streaks being a useful way of predicting the future unless significant evidence were provided.
edit: sorry, misread "one of the" as "the"...
- check bovada.lv for the line
- guess that as the winner
- fin.
(crowdsource ftw)
Streaks from last year's team? Seriously?