Statistics for Hackers
speakerdeck.com
speakerdeck.com
On the other hand, I think there is not enough written about how to do statistics with real, large, complex data sets. I tried to write something like this (https://pavpanchekha.com/blog/stats1.html) but of course the difficulty is that investigating a complex dataset is by definition too complex to really fit into a blog post. In a large dataset, the difficulties are the how and when you make decisions about what to analyze.
It's not complex, but it requires answering scaleability questions. How do you fit the data into memory? How do you even get big data?
I managed to analyze 5.6M Hacker News comments (http://minimaxir.com/2014/10/hn-comments-about-comments/ ) by writing a scraper which dumps everything into a PostgreSQL database (https://github.com/minimaxir/get-all-hacker-news-submissions... ) because it's way too big to fit in memory. Likewise, to visualize the recent NYC taxi data, (http://minimaxir.com/2015/08/nyc-map/ ), I just used BigQuery and saved myself a load of time downloading the data and setting things up.
But in many cases, data is not easily accessible, and not cheap to store and process. I want to switch to an actual big data workflow with Spark, but running the numbers with EC2 storage and processing, it won't be cheap and I may have to set up a Patreon to subsidize things.
On EC2, 6TB of RAM would be 25 r3.8xlarge instances, which is US$70 per hour. If your six-terabyte task is batch processing rather than interactive, that might be cheaper than spending a lot of time optimizing your software for spinning-rust relics.
Contrariwise, if it's an interactive computing task that you can do with AWS Lambda and which you can split into small pieces, 4000 Lambda functions each with 1.5 GiB of RAM, for a total of 6TB, will cost you US$0.01 per 100 milliseconds, or US$0.10 per second. Supposing your overall in-RAM computing task can be done with 10 seconds per Lambda function (a total of 40 000 CPU seconds (11 CPU hours)), can actually be parallelized to that extent, and that 10 seconds is an acceptable response time, then you can do it for US$1. I feel that this must be a mistake, since this is 700 times cheaper than EC2.
(Lambda surely has some total capacity limitations, but I imagine they're quite a bit more than 6TB.)
I am deeply unhappy about the move to such a centralized infrastructure for the internet, but at least this version of centralized infrastructure allows you to run your own code on the centralized servers, unlike the API vision promoted by Facebook and (mostly) Google.
It really is a pity that introductions to statistics spend so much time on analytic approximations and so little on the underlying concepts.
It isn't a hack to present mathematics in an understandable way - it's a pedagogical improvement for introductory works, not a "hack".
This method of presentation makes sense to me. Most statistics classes that I've experienced were taught from the point of view of abstract math. That's certainly one way to do it, but I knew a lot of people for whom that wasn't the optimal presentation strategy. Now that computing is cheap and there is a large audience of people with programming knowledge, I think that teaching statistics through examples of simulation, bootstrapping shuffling and cross validation is a great way to learn things.
Plus, doesn't every hacker start as a script kiddie? Using something others came up with usually sparks the imagination. The next step is usually using the scripts right, followed by using the scripts extensively and finding issues, followed by making your own scripts (or improving the existing ones) to avoid these issues.
This looks pretty hackish to me.
Machine Learning for Hackers http://www.amazon.co.uk/Machine-Learning-Hackers-Drew-Conway...
Design for Hackers http://www.amazon.co.uk/Design-Hackers-Reverse-Engineering-B...
Bayesian Methods for Hackers http://www.amazon.co.uk/Bayesian-Methods-Hackers-Probabilist...
EDIT: I'm not the author, but you can find Bayesian Methods for Hackers (free, released by the author) at the link below. I think it's a great resource for anyone wanting to explore Bayesian methods using Python.
https://github.com/CamDavidsonPilon/Probabilistic-Programmin...
I've seen many, many data scientists from startups write blog posts with skewed data without using such techniques and failing to identify the potential statistical problems.
Any idiot can run a simulation, or compute "something". "Something" is only useful in context.
"If you can write a for loop, you can do statistics!" Ugh. Can nobody read anymore? It has to be a "slide deck"?
Jesus, pick up a book and ask some goddamn questions if you're interested.
Calm down.
In those cases, finding a programmatic hack around it is a very good approach for giving you reasonable results in a shorter timeframe.
http://www.meetup.com/Multithreaded-Data/events/225205209/
I think you're taking things a bit too far here.
[1]: Where poets is a possibly inaccurate metonym for "less mathematically sophisticated students"
* The Elements of Statistical Learning, by Hastie, Tibshirani and Friedman (https://web.stanford.edu/~hastie/local.ftp/Springer/OLD/ESLI...). * Probability Theory: The Logic of Science, by ET Jaynes (http://bayes.wustl.edu/etj/prob/book.pdf)
Best of luck. I can see from your post that you're thinking about performance tuning, I'm assuming you mean of software. That's a nice area - the nice part is that compared to fields like medical genetics, data on performance of software is relatively cheap to get, so a lot of issues about small sample sizes are surmountable.
Bootstrapping doesn't work for very low n (e.g. n=10) because the resamples are not smooth enough and it doesn't always work that well for estimating quantiles. But analytic methods fare pretty poorly in these circumstances too and in any case people are mostly interested in confidence intervals around the mean anyway.
See also here: http://stats.stackexchange.com/questions/172920/bootstrap-me... They give example of non pathologic distribution for which bootstrap doesn't provide good estimate. It's unform(0, theta)
You misunderstand the point about the Cauchy distribution in the answer on Cross Validated. The Cauchy distribution is a degenerate case, mostly interesting as an academic toy because it has infinite variance. Of course that's not going to fare well.
Dependent data can be tricky to deal with, but you can bootstrap such data by removing the dependence, bootstrapping the independent data, and adding the dependence back in. This sounds hard but is usually as easy as running a regression and subtracting/adding the component (x*beta) that leads to the dependency. Alternatively, for timeseries there's window methods.
Of course a short slide deck like the one linked to in this thread is not going to teach you all the finer points of bootstrapping and you can definitely do it wrong. But compared to all the assumptions that frequentist statistics makes to generate confidence intervals and the fact that you need a different method for every different scenario, bootstrapping is about as robust and idiot-proof as it's going to get. "It has its applicability" is beyond selling it short.
Actually, it doesn't. You need a confidence interval, and central limit theorem doesn't give you a confidence interval. It just says that the distribution is close enough to normal at some point. In some cases, it might take a very large n before it become close to normal.
What I wanted to say that you can just blindly apply bootstrap for every distribution. You should carefully check applicability conditions, before using it, or you could easily get nonsensical results.
https://izbicki.me/blog/how-to-create-an-unfair-coin-and-pro...
George Bernard Shaw would refer to this individual as an unreasonable man, but I prefer hacker.