Data Science from Scratch: First Principles with Python
joelgrus.com
joelgrus.com
> As I write this, the latest version of Python is 3.4. At DataSciencester, however, we use old, reliable Python 2.7. Python 3 is not backward-compatible with Python 2, and many important libraries only work well with 2.7. The data science community is still firmly stuck on 2.7, which means we will be, too. Make sure to get that version.
I use the more popular scientific libraries, e.g. numpy, scikit, nltk....and the bigger ones seem to have been ported over to 3.x. A few libs that haven't that come to mind: mechanize and opencv. Has anyone here had success with using 3.x as a data science professional, or is there some massive gaping hole that I'm missing? (I agree that, "Well, this is what the company has been using" is a decent enough excuse to stay on 2.x in most situations)
[1] https://github.com/joelgrus/data-science-from-scratch/blob/m... [2] http://scikit-learn.org/stable/modules/generated/sklearn.clu...
As a side note, do you attend any data events in Seattle? I'm moving there in June after graduation and would love to talk with somebody doing my dream job.
(And I didn't know that until you asked, I'm going to edit the blog post.)
For the combined ebook/print package the code does not work. "You did not meet the criteria for this discount"
[1]http://www.amazon.com/Programming-Collective-Intelligence-Bu...
1. Vector Calculus, Linear Algebra, and Differential Forms A Unified Approach http://www.amazon.com/Vector-Calculus-Linear-Algebra-Differe...
2. OpenIntro Statistics 2nd Edition https://www.openintro.org/stat/textbook.php
Have you posted to DataTau?
If anyone knows the book, can you give a quick overview of how much, math, stats, programming and comp sci. you'd need to read this book? Thank you.
Most of the math is vector space arithmetic. There are a few sections that use matrix multiplication. The probability and stats is stuff like understanding probability distributions and Bayes's Theorem. (It's all covered in the book, but you'd need to be comfortable picking it up and using it.)
In terms of programming, not much. Someone who's never programmed before would probably have a tough time, but the goal is that someone who is bright and hardworking and who can write fairly simple Python programs should not have a problem. Very little CS background required. Maybe basic data structures like list vs dict and so on.
I'm thinking about it in terms of running computation in production environments where you may be constrained by available compute resources or budget. Some people have an intuitive grasp of cpu/memory/bandwidth and can do performance tuning as necessary, but those who don't can find themselves in situations where they waste a lot of resources, such as running a million parallel jobs that each have less than 1 second of CPU time, getting stuck after failing to request or provision nodes with sufficient memory, or performing unnecessary reads and writes.
http://shop.oreilly.com/product/0636920030157.do
It's more focused on how to analyze existing biological data with the shell, R, and how to use git.
Personally, I've rarely seen advanced machine learning being used outside of genome-wide association studies, and even there most people just use PLINK's logistic regression without understanding what's being done and call it a day.
Another really good book on how to understand statistics is Motulsky's Intuitive Biostatistics - it introduces all common "tests" and methodologies people working in the life sciences use, but without the formulas (you use R for that anyway). It's more about the caveats of each test, in which situation you'd use it, what can go wrong, how to interpret the results etc., all written in a very lively style.
I can't help but wonder about those recommender systems. With so little material on statistics I have to assume it's only about observational data, which is the best way to make the millionth+1 useless recommendation engine.
And why is it that the "data science" books never discuss DoE?
To be fair, ml tends to focus very heavily on prediction, not inference / interpretation of betas. In many tree models how to even understand coefs is an open question.
Most of the big ML books are heavily Bayesians, and these subjects are less discussed (though IIRC Gelman's book has a "Bayesian ANOVA"). Even Elements of Statistical Learning, which is very frequentist in its approach, only references ANOVA in passing. Do you have any book to recommend about these fundamentals?
intermediate level, covers some blocking IIRC: _Statistics for Experimenters_ by Box et al
advanced: I thought quite good, but classmates did not universally love. Unfortunately does not come with case studies or R code to run them; I have a bunch but (very unfortunately) printed instead of computerized and, in any case, probably copyrighted by my professors. _Experiments: Planning, Analysis, Operation_ by Wu and Hamada. The math is not complex but can be involved for various types of blocking designs.
How do you feel about the Bayesian approach to these questions? (cf. Gelman's Bayesian Data Analysis)