Every NFL play for the past 10 years in CSV format
advancednflstats.com
advancednflstats.com
20090201_PIT@ARI,2,30,18,ARI,PIT,1,1,1,(:18) (Shotgun) K.Warner pass short middle intended for A.Boldin INTERCEPTED by J.Harrison at PIT 0. J.Harrison for 100 yards TOUCHDOWN. Super Bowl Record longest interception return yards. Penalty on ARZ-E.Brown Face Mask (15 Yards) declined. The Replay Assistant challenged the runner broke the plane ruling and the play was Upheld.,7,10,2008
They forgot: for(i=0;i<92;i++){yell('edw519','GO!')}
Seriously, I had plans for the next 4 days, but I just scrapped them. Funny how jazzed I get when it's data that I can really relate to...
I've already structured my data warehouse and started the loads. (I'll probably need a whole day just to parse the text in Field 10.) Then I'm going to build a Business Intelligence system on top of it. I will finally have the proof I need that I, not the offensive coordinator, should be texting each play to Coach Tomlin.
See you guys on Monday.
EDIT: OK, I'm back, but not for long. I'm having way too much fun with this...
fleaflicker: Cool website & domain name. Thanks for the tips. I expect shortcomings in the data, but it looks like it's in a lot better shape than the usual free form enterprise quality/vendor/customer comments I usually have to parse. We'll see...
MattSayer & sjs382: I don't plan to do any analysis. I prefer to build an app that enables others to do their own analyses, answering questions that nobody else is asking. Like "Which Steeler makes the most tackles on opposing runs or more than 5 yards when it's 3rd down and longer than 8 yards to go, the temperature is below 38, and edw519 is twirling his Terrible Towel clockwise?"
jerf: Nice thought. I've spent years trying to earn enough money to buy the Pittsburgh Steelers just to fire the slackers and fumblers and win the Super Bowl every year. Maybe I should just take an easier route and solve that problem like any self-respecting hacker should: with data & logic. No Steeler game this weekend; I may have found my destiny </sarcasm>
First thing that I noticed was that the Game ID matched CBS's website's URLs: 20090201_PIT@ARI == http://www.cbssports.com/nfl/gametracker/playbyplay/NFL_2009...
Also, I went to compare this PBP to both ESPN and CBS and found that both have the exact PBP data, which is interesting because it seems that they got this data directly from the NFL (or from the same source, at least). I guess this makes sense, but it's something I hadn't considered.
For reference, ESPN's PBP for the same game: http://espn.go.com/nfl/playbyplay?gameId=290201022&perio...
Overall this is neat but it's hard to find real life context within this data. Was the QB pressured, was a coverage blown, was there a pre-snap audible or motion or change by the defense, what was the formation, how much sleep did the players get the night before, etc etc.
For example, first initial plus last name does does not uniquely identify a player. You'll need accurate roster data first, and even then there are clashes.
We store play data by its structured components (players involved, play type, player roles, etc) and then derive the text description. This allows us to reassemble pbp data from different pro games to show a "feed" for your fantasy team.
Baseball has a smaller set of play outcomes/transitions so its easier to model this way. As your example from the Steelers Super Bowl shows, football plays can be very complex.
Which means you can treat it a bit like a text mining program. NASA had a text mining contest in 2007 as part of the SIAM conference on data mining which was really similar - instead of football plays it was textual descriptions of aeronautics incident reports and their classification. There were several papers that came out of that (I was with a group that did one of them, using an approximate nonnegative matrix classification approach - got beat out by some ensemble approaches).
Anyway - if you'd like to do something with unstructured football play descriptions, text mining might be able to empower you to some extent without going through a full manual analysis, and those papers could be a good starting point. I think some of them ended up in a volume titled _Survey of Text Mining II_.
But for high-dimensional text problems that're pure classification, I tend to rely simply on 1NN classifiers (against a single centroid of training data of a target category, of which there tend to be many). I've spent a lot of time with NMF, for its potential as an incredibly interesting data-exploration tool ("There's a pronoun cluster! There's a Spanish cluster! There's a 404 Error axis!") or low-dimension projection step. I've even spent a good amount of time on implementing the algorithm in a number of memory-efficient ways.
Could you expand a bit on how you used NMF for these problems in practice (similar to how a sparse autoencoder captures reduced-dimensional features en route to supervised learning), or how others used ensemble methods?
The main idea, though, is to generate a term-by-document matrix (count words, maybe throw out stopwords, normalize counts), then do Math to factor your matrix (approximately) into two: term-by-feature and feature-by-document. When you want to classify a new document, you can use its contents (more terms) to calculate a feature vector.
(The math seems to typically involve random initialization followed by iterative improvements. Other work in the field discusses the specifics.)
The matricies are "nonnegative" because, conceptually, features are a _positive_ thing, and you can't say that a certain term makes something less a member of a feature cluster (only more).
The tricky part is figuring out how to map features to things which are semantically interesting to your application, and I don't want to comment too much on the state of that because it's been five years and I honestly forgot what exactly we did there, and it was all done in Matlab (which I'd never used before), and there's probably more recent work in the field. But if you fiddle with it manually, you can come up with your matrices and essentially have a nice little classifier.
Thanks!
http://www.advancednflstats.com/2010/01/expected-points-ep-a...
I've been trying to work at the college football level with this same strategy but I'm still trying to figure out how its calculated. It seems trivial but it takes a lot of data organizing.
Though perhaps not as open...
It seems this might result in bugs, as in the Oct 20, 2002 game between Dallas and Arizona. In the third quarter, with a score of Arizona 6 - Dallas 0, Dallas scored a touchdown (row 13900) but "aborted" the extra point (row 13901). The 6 points for the Cowboys are not recorded in the data.
The game eventually went to overtime, with the Cardinals kicking a winning field goal in OT for a final score of Arizona 9 - Dallas 6, but the data here records it as Arizona 6 - Dallas 0.
http://www.advancednflstats.com/2007/02/contact.html
Of particular interest:
Where did you get your data?
Most of my team data comes from open online sources such as espn.com, nfl.com, myway.com, and yahoo.com. It's easy for anyone to grab whatever they're interested in from those sites.
My play-by-play data comes from a source that's not publicly available, and at this time I regret that I cannot share it. However, I am working hard to develop a way to spread the wealth. One of my biggest goals is to help create a larger, more open, and more collaborative community for football research.
----
There's no real terms of service so I'm curious as to the constraints in using this for commercial purposes. I most definitely want to use this for teaching purposes (how to text-mine, how to build a web app from data, etc) but want to know what terms the data can be redistributed.
http://blogs.trb.com/sports/custom/business/blog/2009/04/cbs...
However, you do have to get the data, and unauthorized access of computers (which constitutes trespassing) can be a legal gray area. I'd love to hear a lawyer weigh in on the legality of scraping the data directly from espn.com.
http://nfl-query.herokuapp.com/
The basic syntax is [stats] [conditions] : [row] / [column].
There's some autocompletion to try to make it possible to discover what is accepted.
Examples:
passing yards : team / season
first downs / first down attempts : down / distance
rushing yards min 100 rushing yards : player, game, quarter
rushing yards / carries min 200 carries : player
One of the biggest problems is that it's currently way too easy to shoot yourself in the foot by making a really slow query.
It seems to me that the NFL would want to have exclusive rights to distribute this data and charge people a fee for access to it. Clearly I'm no expert in these legal affairs though.
On the other hand, a live performance is generally not protected from copyright, so if you attend a live game to collect the data, you may be in the clear.
The data isn't owned by the NFL, but all recordings of the games are, and so any data obtained by watching recordings of the games could potentially be controlled by the NFL.
I wouldn't want to put a large bet on where exactly those lines are drawn, though.
Edit: I should note that even http://gdx.mlb.com/components/game/mlb/ contents are governed by http://gdx.mlb.com/components/copyright.txt which doesn't allow commercial use.
IANAL, but I worked at ESPN and founded Fanvibe (YC S'10), and worked quite a bit with the leagues and lawyers on rights-related topics.
IANALTINALA.
I can't even give you a list of football fixtures coming up this weekend without breaking copyright law.
In his early Stanford days, Bill Walsh had already cracked the code on how un-random football coaches (and almost all people) are. From "Controlling the Ball with the Passing Game":
"We know that if they don't blitz one down, they're going to blitz the next down. Automatically. When you get down in there, every other play. They'll seldom blitz twice in a row, but they'll blitz every other down. If we go a series where there haven't been blitzes on the first two downs, here comes the safety blitz on third down."
Most NFL offenses tend to alternate rather than randomize. Walsh knew defenses were just as predictable decades ago.
[1] - https://github.com/BurntSushi/nflgame
[2] - http://blog.burntsushi.net/nfl-live-statistics-with-python
SELECT off, COUNT(off) AS count
FROM [nfl.2012reg]
WHERE description CONTAINS "INTERCEPTED"
GROUP BY off
ORDER BY count DESC
(This counts the number of plays that resulted in an interception by the team that threw the interception, sorted from most to fewest INTs)http://www.football-data.co.uk/downloadm.php
Tons of European leagues, going back to 1993 in some cases.
Here's some sites that give detailed stats and match reports:
Man City use their petro-dollars to open up Opta Sports (detailed match stats) to all: http://www.mcfc.co.uk/the-club/mcfc-analytics
Someone needs to compile stats equivalent to these NFL ones for european football! Hmmmm...
http://www.advancednflstats.com/2009/09/4th-down-study-part-...
The tl;dr version can be found at:
http://www.advancednflstats.com/2010/05/4th-down-briefs.html
The conclusion is that teams should go for it on 4th down much more often than they currently do.
He also has a calculator where you can get the exact values:
http://www.math.toronto.edu/mpugh/Teaching/Sci199_03/Footbal...
I also believe Belichek/Adams have funded some football economics research.
Sample description field: > 20020905_SF@NYG,1,59,20,NYG,SF,3,11,81,(14:20) (Shotgun) K.Collins pass intended for T.Barber INTERCEPTED by T.Parrish (M.Rumph) at NYG 29. T.Parrish to NYG 23 for 6 yards (T.Barber).,0,0,2002
In the comments section of the OP, someone posted this sample Excel function:
=IF(ISNUMBER(SEARCH("right tackle",J2)),"rush",IF(ISNUMBER(SEARCH("right
guard",J2)),"rush",IF(ISNUMBER(SEARCH("left
guard",J2)),"rush",IF(ISNUMBER(SEARCH("up the
middle",J2)),"rush",IF(ISNUMBER(SEARCH("left
tackle",J2)),"rush",IF(ISNUMBER(SEARCH("left
end",J2)),"rush",IF(ISNUMBER(SEARCH("right
end",J2)),"rush",IF(ISNUMBER(SEARCH("pass",J2)),"pass",IF(ISNUMBER(SEARCH("kneel",J2)),"kneel",IF(ISNUMBER(SEARCH("punt",J2)),"punt",IF(ISNUMBER(SEARCH("kicks",J2)),"kickoff",IF(ISNUMBER(SEARCH("extra
point",J2)),"extrapoint",IF(ISNUMBER(SEARCH("sacked",J2)),"sack",IF(ISNUMBER(SEARCH("PENALTY",J2)),"penalty",IF(ISNUMBER(SEARCH("field
goal",J2)),"fieldgoal",IF(ISNUMBER(SEARCH("FUMBLES",J2)),"fumble",IF(ISNUMBER(SEARCH("spiked",J2)),"spike",IF(ISNUMBER(SEARCH("scrambles",J2)),"rush","rush"))))))))))))))))))
Dear god, at what point do people finally realize that it's worth learning some simple scripting to work with text files?At any rate, what would you recommend most to accomplish the task? I'm learning Python and know R a bit, so I was just wondering how I was going to about combing through the data.
Python and Ruby would also allow for more elegant-looking -- i.e. more maintainable -- functions to handle that field.
Other than cool data visualization stuff, the obvious implication is the potential to devise a profitable system to pick games against the spread. The guys at Football Outsiders have done a decent job at it and made a proprietary algorithm that picked games at 58% this year ( which is over the threshold you need to be profitable in Vegas ). But even those guys are still having some trouble getting access to and aggregating the data in a usable format.
I really want to sit down and start playing around with some of this data so I appreciate you putting this together for everyone. The NFL needs an open source API and this is definitely a step in the right direction.
btw, pro-football-reference has pbp data now too, and it probably goes back a lot further, but I think they discourage mass scraping of their site.
Like someone else said, it's about match-ups.
Semi-related : http://profootballtalk.nbcsports.com/2013/01/03/polian-think...
http://www.nfl.com/news/story/0ap1000000121582/article/bill-...
When all was said and done it worked, but made some pretty crappy draft picks! I should find that code....
Head Coach - Offensive Coordinator - Defensive Coordinator - Formation - Play