How Coaches and the NYT 4th Down Bot Compare
nytimes.com
nytimes.com
Also (and to their point), there is a ton of variation between offenses, defenses, punters, running backs, etc. Using this as a rubric is a nice idea for a rule of thumb for an armchair quarterback, but I'd strongly disagree with using it to accurately criticize any decision within a couple points of expected value.
Until you get a coach's motivations more in line with the motivations of the team, you won't get a coach that is as aggressive as these computer simulations suggest they should be.
However, the obvious counter to your question is how many coaches really go for it frequently on fourth downs? Most coaches just follow the conventional wisdom unless they are already in a desperate situation like Ron Riviera or they have built up a lot of good will like Bill Belichick.
He hardly made a controversial decision.
Since coaches are conservative, they're likely to go for fourth down conversions only in cases when they are likely to succeed (due to defensive mismatches, etc).
Hence the model may over-estimate the benefit of going for it on fourth down.
Since the data from this is from previous plays in the last 10 or so seasons it is flawed. There probably isn't very little data of a team going for 4th and 1 and that deep in their own territory. The only time this has probably happened is in extraordinary circumstances - specifically in a game-winning drive scenario with time expiring, where the defense is in 'prevent'. In these cases it is highly successful in that situation but you don't have the time to get the score.
Meanwhile, on fourth down, the defense is amped up and ready to roll. You'll usually see a strong blitz against a sputtering offense, and the percentage of plays that actually get a yard is going to be lower.
Not only that, going three-and-out means that the offense is unlikely to do any better going from the 10 to the 20 yard line. Better to punt the ball, trust your defense, and have your offensive coordinator fix whatever is going wrong.
That may sound like a lot, but I'd imagine the actual point difference is on the order of 2 points.
The more interesting thing to consider would be whether taking a safety there is ever a positive.
It seems though he gave up some amount of EV by giving up the 2 points to the Broncos, thus a FG later on would only tie the game instead of winning it. But I assumed this was offset by the risk of having to punt and the Broncos end up getting a FG where then they'd be up by 4 points.
I don't know if Belichick was right to do this, but it paid off for him which is why everyone pats him on the back.
Contrast this with the 2009 Colts/Pats game where he went for it on 4th and 2 on his own 28 and didn't make it.
Going for it on 4th and 2 in the '09 game was the right decision, and few knowledgeable commentators disparaged Belichick for it. Manning was dissecting the Pats defense at that point in the game: He had a much greater chance of scoring a touchdown than the Colts defense had of stopping the conversion.
Now the playcall...
This bot also requires both coaches to follow it's recommendations, because otherwise the comparison between what they call "point values" of the two teams becomes meaningless.
I've thought about it some since then and I'm also not convinced that expected points is a good metric. That will maximize the expected point differential per season but I am not convinced (although I could be wrong... I haven't put pencil to paper to make my thoughts rigorous) that there are enough possessions in the game on average to make this metric useful even in the first half.
http://www.advancednflstats.com/2013/11/momentum-1.html
That jibes with my own anecdotal experience watching a lot of football, despite how often announcers like to cite it as a factor.
Also, looking at the overall pattern of the data, my gut tells me that the optimal rubric is somewhere in-between the 'bot and the meatcoach. There are some strange holes as well - at 4th&3, you should go for if you're on your own 8 or 9, but punt if you're on your own 7 or 10?
But the holes at the your own end you mentioned, that makes no sense at all.
Still, there is probably a case to be made that coaches should be a bit more aggressive, but that edge could be small enough that other factors (like perceived incompetence) dominate it.
This seems pretty loose. I'm not trying to be pedantic, but the theoretic flaws in the data are such that you would want a pretty tight logic in your model to have any level of comfort with it. The proportion of the state-space that is unobserved seems ~large and critically relevant.
You can't reliably predict the future with statistics alone. If you could, then technical analysis of the stock market would be the most successful investing strategy. It is not.
Likewise, a lot of folks seem to think that taking an historical statistical trend and extending it into the future is a scientific approach to football coaching. It's not.
Scientists use statistics to understand evidence, but they formulate theories in terms of root causes. In other words it's not enough to know that dropping a hammer has a high correlation with acceleration of 10 meters per second squared. If you want to predict how fast other objects in the future will fall, you need a theory that describes why and how anything falls.
It doesn't look to me like the 4th down bot embodies or implements any cohesive theory of football success. It just finds trends and dots the line into the future.
An interception is better or worse than turning it over on downs by the field position delta. If it is returned past the line of scrimmage, then it's worse. If it is not returned, it's better. A deep interception with no return is like a punt in terms of impact on the game.