I'm not a statistician, so I can't really speak to how 'valid' the analysis is, but I'd be curious to see how it does in different tests--unless I'm misinterpreting, the biggest check so far was done on 2012, which is exactly what was used to train it. It would be interesting to see what happens if you train with half of 2012 and test the second half. Or check 2011 (do you predict an end-of-the-year collapse, allowing my Cardinals to sneak in again? ;) ).