How to create NBA shot charts in Python
savvastjortjoglou.com
savvastjortjoglou.com
...or if you're in the beta, here: https://www.statmuse.com/nba/search?q=michael%20jordan%20sho...
Looks like they're translating natural-language into some sort of query language. Are there any well-known methods for doing this kind of thing?
[1]http://www.vldb.org/pvldb/vol8/p73-li.pdf
Edit:
Since this got brought up again I've been searching around a little more and found c-phrase, which looks similar.
https://code.google.com/p/c-phrase/ https://www.youtube.com/watch?v=fWio8bHq4wQ
The heat maps are much less useful IMO, because the colors are poorly chosen and the data generalization/binning methods seem kind of arbitrary. Until the shot count gets up into the tens of thousands or more, just show all the data. (If e.g. aggregating all the shots from the whole league, then some kind of binning would become necessary.)
The marginal histograms showing density by x/y coordinates on the court are essentially useless in my opinion. Dramatically more interesting would be marginal histograms related to angle and distance in terms of polar coordinates centered at the basket (it might be necessary to ignore the angles for positions very close to the basket, where angle is kind of irrelevant). To make them even more informative, since the 3 point line isn’t a perfect semicircle, make marginal distributions (in terms of angle and distance) of separate categories of 2 point and 3 point shots, and stack them. Or even three categories showing dunks/layups, 2 point jump shots, and 3 point jump shots.
Just a couple of questions: What color maps would you use for those kde plots?
I know seaborn by default uses the Freedman-Diaconis rule to create bins for the hexbin plot. But, what suggestions do you have for binning?
http://www.basketball-reference.com/players/h/hardeja01/shoo...
http://www.basketball-reference.com/play-index/plus/shooting...
if so, would you mind if I asked you a somewhat technical question?
Is there some sort of advanced query function?
-------------------------------
Here's kind of what I'd like to be able to do
http://fivethirtyeight.com/features/no-team-can-beat-the-dra...
the first chart in there sort of fits a line from draft position to career AV (7 paragraphs in)
I'm curious if you can hold for certain factors and move that line around
specifically I'm curious if the line changes if you adjust for winning percentage of the player's college team
ie if you did a curve players from schools w/ winning percentages >.750, .750 - .500, .500 - .250, .250 - 0, would that produce 4 noticeably different curves?
------------------------------------
it appears that all the data necessary to do that is contained between NFL-reference and CFB-reference
the way I would think to do that is to scrape the data off the site and put it into an excel sheet to work with
is that the best way to do that?
or is there a function within the site that I can work with such that I don't have to scrape the data?
-------------------------
thanks, sorry if this sort of question was outside the bounds of this message board
http://www.pro-football-reference.com/play-index/draft-finde...
www.datasciencemasters.org -- comprehensive list of things to learn / explore
www.dataquest.io -- (disclaimer: I'm involved with the company), teaches data science in the browser via projects