Open Football Data
openfootball.github.io
openfootball.github.io
For details and advanced analytics though, this one is much better: https://github.com/soccermetrics/soccermetrics-client-py
Lots of numbers to crunch there!
This is all very nice but it would be nicer if there was some sort of cheap software that amateur teams could use to gather and then analyze their own data. There's a massive market out there for this sort of thing, the football world is very conservative and tends to move slowly.
I'm the founder of Soccermetrics and the creator of the Soccermetrics API. Thanks for the attention.
I've had the site for a little over five years now. I had been working on data models and algorithms for analysis of soccer matches and thought I would have a go at creating a company out of it. I even applied to YC which was a bit of a laugh in retrospect :) Right now I have another job that pays the bills but there are a few projects I do on the side, the API being one of them.
The API is the latest iteration of my data models exposed to the world as well as my attempt to build as close to a REST API as I could. I don't claim perfection and I'm sure others will have their opinion on it, which I welcome.
I wrote the Python client which is as you say a wrapper over the API, which serves well as a starting point. If you have ideas on how to extend it, please fork and contribute.
There are a few software tools out there that do what you wish. Statzpack is one, SportyBird is another. I have my doubts about how big the market really is for this kind of service, but everyone is very early in this space.
My family (everyone except me, lol) sort-of runs an amateur football club, so I have some second-hand knowledge of that world (at least in Italy). They tell me that systematic, professional and data-driven approaches are incredibly scarce but very effective. It's a system that still runs on personal networks and a lot of outdated knowledge and "magic", not unlike the baseball world described in Moneyball [1]. As you said it's early days, but still, most coaches under 40 now bring a tablet with them on the bench, and not to take funny pictures.
[1] http://www.amazon.co.uk/gp/product/0393324818/ref=as_li_ss_t...
By the way, thanks for the pull requests on the client. Developing the client and the API backend has been a learning experience at every step, so I appreciate contributions from experienced developers.
Already thinking about the apps that will use this! Thank you.
The more the better!
Something around the tune of $25k a year. Anyone actually paying for this now and can provide pricing?
https://github.com/opensport/american-football.db
The best public football repository I am aware of is this, though:
http://www.advancedfootballanalytics.com/2010/04/play-by-pla...
I thought I'd just squeeze in a few words about nflgame/nfldb. Both offer access to the same stuff: play-by-play data back to 2009. Both can be used with live games so that they are updated in real time (well, at least as frequently as NFL.com).
nflgame is responsible for pulling the JSON data and provides some rudimentary searching features. But it's slow.
nfldb stores all this data for you in a relational database. It comes with a script that updates the database while games are playing so that you can get access to live data. (It will even migrate the database for you if I've made any changes to the schema.)
Here's a quick example that shows how to get all of Julian Edelman's touchdown plays from last season:
import nfldb
db = nfldb.connect()
q = nfldb.Query(db)
q.game(season_year=2013, season_type='Regular')
q.player(full_name='Julian Edelman').play(offense_tds=1)
for g in q.as_plays():
print g
Easy as pie!There's an extensive wiki (almost 20,000 words) with tons of examples and explanation: https://github.com/BurntSushi/nfldb/wiki
Other features: aggregating data, player meta data (college, height, weight, etc.) and fuzzy player name matching.
The data format seems to be a custom text format which admittedly I could be wrong about. Is it possible to use TSV or CSV instead since it would be infinitely more useful since it could be directly imported into relational databases, Excel, etc.
The short answer is no. I've searched long and hard, high and low, for free (beer) horse racing databases for UK/IRE and Australia. To a lesser extent I've searched for HK, FR and GER data. I'm yet to find anything that is comprehensive and no cost.
There's a couple that I do use for UK/IRE racing which cost in the region of £35-£45 per month for access. Betwise/Smartform provides an historical database in MySQL, and daily race card/results updates. UKHorseRacing.co.uk provides CVS files with historical race data, their ratings and race results. I take these CVS files, combine them into a SQLite database and interrogate with R.
A slightly longer answer is, sort of. The Betfair API is currently open access for non-commercial and low volume use (as far as I'm aware). This will allow you to retrieve basic racing data - the cards before that race with horse name, jockey, barrier etc and the race results post-race including the Betfair Starting Price. After interrogating the API, you'll need to obviously compile the data into your own database. A bit of work, but feasible. Betfair has a developer programme and their are API bindings available in a number of different languages. I use R (R package developed by Betwise mentioned above), but I know Python is available. One caveat to mention is that Betfair are upgrading their API, so this will obviously have an impact on existing programs using the old one.
If anyone else has additional information or could point me in the direction of something else "free" I'd appreciate it as well.
It is free but you need an active account with them to download the CSV files.
http://thorotrends.com/news-and-views/50-blog/117-release-th...
http://www.anddownthestretchtheycome.com/2012/1/14/2706205/t...
http://www.paulickreport.com/news/ray-s-paddock/free-our-sta...
## GK / Goalkeepers
Kawashima|Eiji Kawashima, 20 Mar 1983
Nishikawa|Shusaku Nishikawa, 18 Jun 1986
Gonda|Shūichi Gonda, 3 Mar 1989
## DF / Defenders
Inoha|Masahiko Inoha, 28 Aug 1985
G. Sakai|Gōtoku Sakai, 14 Mar 1991
Nagatomo|Yuto Nagatomo, 12 Sep 1986
Uchida|Atsuto Uchida, 27 Mar 1988
Konno|Yasuyuki Konno, 25 Jan 1983
Kurihara|Yuzo Kurihara, 18 Sep 1983
H. Sakai|Hiroki Sakai, 12 Apr 1990
Yoshida|Maya Yoshida, 24 Aug 1988
Masato Morishige, 21 May 1987 ## Japan F.C. Tokyo
Comments as a double-hash, key fields are either player last name or occasionally first initial-space-last name, then three different delimiters of pipe, then comma, then tab. Choosing either a consistently delimited format or a more verbose JSON/YAML structure with clear metadata would seem to be a better approach.[1] https://github.com/openfootball/players/blob/master/asia/jp-...
how often is the feed updated?
Kickdex is also pretty awesome, they use the Opta data to produce real time indices for teams and players.
For copyright protection to apply, the database must
have originality in the selection or arrangement of
the contents and for database right to apply, there
must have been a substantial investment in obtaining,
verifying or presenting its contents. It is possible
that a database will satisfy both these requirements
so that both copyright and database right apply.
They would have a "database right" if they had placed a person at each match to gather the data and verify it, as that is a substantial investment.How they originally acquired the data is important and shouldn't be presumed.
However that doesn't stop you from implementing your own database and re-acquiring the facts in some trivial way. Just bear in mind that accessing historical data may breach someone else's database right.
Database rights are usually proven by fake data inserted into the database to catch people copying it.
For example you could argue that the Rare Record Price Guide ( http://www.rarerecordpriceguide.com/ ) is just a collection of facts, and decide to copy it... but you'll discover when sued, that a few of the bands in the guide are fictional and designed to demonstrate that the database is theirs, and that it's not trivial to acquire and verify the data.
So, for the sake of argument, if the dataset had no fake data then it would be OK? Or would they still need to demonstrate "substantial investment", no matter the state of the data?
If the latter, then that gets weird quick. How many lines of code is considered substantial? How many hours hunched over a microfiche machine? It sounds like it would ultimately depend on the skill of your lawyer.
If you have enough data sources you could theoretically recreate a play by play from all of them and have a data set that would be difficult to prove was stolen from someplace in particular. I say theoretically because (at least with college football) you are often not given enough information to recreate the game (simple example would be how long a play took to execute to determine drive possession time), so often you are left using a best guess method.
This is an open source schema for storing data, why not re-acquire the data from a fresh source and make that open source too? This avoids pulling it from an existing and potentially protected source.
You could have members of the public individually enter historical scores, and each one provide proof of the score (e.g. a photo of a result in a newspaper or a photo of the matchday guide).
You could verify correctness of that acquired data by comparing to a few known data sources (even if they were protected). So long as you were close enough in fact to not alter history it was probably right, and correctable in the future (editable like a wiki).
If you use one of the existing datasources you'll find yourself with a lawsuit if you reach any reasonable size.
Details of the case law here: http://www.out-law.com/page-392 And of the court case: http://www.out-law.com/page-5055
The lawsuit backfired completely. The BHB had wanted to charge newspapers for publishing horse racing fixtures, but it inspired the newspapers to turn around and question why the BHB wasn't paying them for devoting pages to the sport.
(Looks like the case also involved football fixtures as well)
There's been quite a bit of legal back and forth
- http://www.football-dataco.com/ - http://www.bbc.co.uk/news/business-17218968 - http://www.twobirds.com/en/news/articles/2012/football-datac...
That practice ended with the March 2012 ruling in the European Court of Justice [1] that neither rights subsist in relation to fixture lists.
However, in a separate ruling [2] Football DataCo successfully argued they did have a database right over live data concerning matches (e.g. goals, goalscores, cards, etc.)
[1] http://curia.europa.eu/juris/document/document.jsf?docid=119...
The license on the data is a pretty permissive one, simply requiring attribution of the data to the Retrosheet project. Software to process Retrosheet files is available, under the GPL:
http://www.seanlahman.com/baseball-archive/statistics/
Then of course MLB has a bunch of data here, mainly the PitchF/X data since 2008 is gathered from here.