Best Data Science Books According to the Experts
builtin.com
builtin.com
* Vectors, Matrices, and Least Squares — IMO the best beginner-friendly and applications-focused intro to (or review of) linear algebra. Covers a ton of fundamental ground while keeping things consistent and concise. Lots of exercises and a Julia supplement book. (http://vmls-book.stanford.edu/)
* Mathematics for Machine Learning — good coverage of the most important math concepts relevant to ML (https://mml-book.github.io/)
* Forecasting: Principles and Practice — best overall resource on forecasting that I know of; R focus. One of a zillion great R/data science books, virtually all of which are open and well-written. (https://otexts.com/fpp2)
* Dive Into Deep Learning — can’t personally vouch for this one but it looks comprehensive; numpy/PyTorch/TF focus (https://d2l.ai/)
* Speech and Language Processing — clear introduction to all things NLP, nice flow (https://web.stanford.edu/~jurafsky/slp3/)
I tend to not be a cover-to-cover reader, so I usually deep dive into a single topic for a while (e.g. forecasting, information retrieval) and read papers/tutorials/chapters related to that topic and the math concepts related to it.
PS I feel like impostor syndrome is so common among data scientists because there is so much material that feels like “must know”. Don’t feel like you need to memorize thousands of textbook pages to be effective, and you could spend a lifetime mastering any of these individual subjects. JIT learning is a great skill to have.
If I had to choose 1 book as the ML bible, then it would be Murphy's (contrasted against Bishop and ESL) for the following reasons:
1. It uses CS jargon. (Bishop's book while great, uses Math/physics notation/jargon which add a barrier to entry)
2. It is more up to date and comprehensive (It covers everything from probabilistic models, traditional models, to neural networks all the way up to 2015 or so, unlike say ESL which is more introductory)
3. Everything in deep learning past 2015, is better learnt through papers/video lectures than any book. (A lot of it is intuition and not truths. There is a certain authority to books that belies our lack of understanding of NNs. Opinions on popular operations such as Dropout, Batch Norm, saliency maps have changed drastically over the last few years)
4. My ML professor used it for our upper grad level ML course and I came out very satisfied. (nothing quite like personal validation). I have read ESL and found it to be better as reading for an intro to ML course. I tried reading Bishop, and didn't like it :| )
Just googling kevin Murphy ML gives me a lot of ads and pages about hair product.
1. https://www.cs.ubc.ca/~murphyk/MLbook/errata.html
2. https://www.cs.ubc.ca/~murphyk/MLbook/pml-print3-ch19.pdf
https://mitpress.mit.edu/books/machine-learning-second-editi...
I will probably revisit it entirely after that.
Then you will want to consult a text book in your work domain (e.g. introduction to speech & language OR statistical natural language processing for the domain of natural language processing).
And finally, you will want either a book, or free online Web resources/tutorial videos that show you how to do things in practice, given a particular programming language and tool-set (e.g. Python + TensorFlow, Java + DeepLearing4J).
This recipe of Theory + Application + Practice/Tools should get you there.
Introduction to Statistical Learning is also available for free online:
http://faculty.marshall.usc.edu/gareth-james/ISL/
Although I only read a few chapters from that book, I really like it (but I would have preferred a python version of the book).
Personally, if you have to pick three books from the list, ypu can start with these three options.
It's an excellent 'zero-to-hero' text for understanding deep neural networks, some common architectures, and the code (and theory) to get them to work.
One thing missing is how to prepare data for deep learning -- but that's just standard ETL you learn elsewhere.
(Edit: I was wrong about affiliate links. Not deleting to publically self-shame.)
The real data science would be producing articles like this automatically, and with good SEO, to drive revenue.
The list is OK. I've studied ISL (and some of ESL). A friend really enjoyed Think Stats. Charles Wheelan's book is in the same vein as How To Lie With Statistics (Darrell Huff?) but in greater depth.
I started with R because that's what my team was mostly using when I got into this area. Hadley Wickham's free books are good too.
The Elements of Statistical Learning (ESL) [2] is also a very reputable source.
They also make a good pair: PRML has a Bayesian approach while ESL is traditionally frequentist.
[1] http://users.isr.ist.utl.pt/~wurmd/Livros/school/Bishop%20-%...
https://github.com/melling/data-science-from-scratch-swift/b...
I haven’t read it yet but the Think Stats book is available for free online:
https://henrikwarne.com/2019/07/27/book-review-designing-dat...
- Goodfellow et al. Deep learning. MIT press, 2016. https://www.deeplearningbook.org/
And there is the book about DL with Python published by Manning in 2017, but not the book about DL with PyTorch, published by Manning in 2020?
- Stevens et al. Deep Learning with PyTorch. Manning Publications, 2020. https://www.manning.com/books/deep-learning-with-pytorch
https://pytorch.org/assets/deep-learning/Deep-Learning-with-...
The "Bayesian Inference and Machine Learning" track gives a nice foundation to anything with "log loss". After that even k-means won't be an ad-hoc algorithm.
It's not focused on data science specifically. But it's short, sweet, and just might leave you better prepared to tell whether the emperor is wearing actual clothes, or if it's just an artfully arranged assemblage of TensorFlow operators.
Got any other companion book recos?
I'm particularly interested in data sets that I could play with in Tableau, Google Data Studio, and Excel.
https://news.ycombinator.com/item?id=23520545
Has quite a few of the books mentioned in the post and the comments here.
I have plenty of textbooks in my reading list already.
If you don't know Linear Algebra; the "done right" book is absolutely not done right for people that work with data in the real world: Strang is infinitely better, and Trefethan and Bau and Gollub if you need to go deep.
As someone pointed out below: Kevin Murphy's ML book is actually very good for teaching you how the things work. What's more, you can download python or Matlab for every algorithm in his book (as you can for Bishop at this point).
Nobody needs the Deep Learning books listed. In spite of the hype, it's just not that important compared to naive bayes, LDA and GBM; most people don't have access to the hardware or data sets that make DL useful, and those who do probably studied DL in grad school.
Gads; what trash! Those experts: only one of them appears to actually be an expert; pedants and post-docs don't count.
Herman:
- ISL
- Hands on
- Chollet
- spark+desk
Miller:
- Grus
- thinks stats
- la done right
- bishop
- data intensivePeople good at (1) and bad at (2) write "PhD code" that may or may not be right but you can't tell because it's too disorganized. People good at (2) but bad at (1) get fine-ish looking numbers out of their good looking code but you can't tell whether it's right because they may have ignored or misunderstood fundamental assumptions and correctness of the underlying methods.
There are also seemingly tens of thousands of people on the market who have little experience in either but have adapted projects from examples online into their Github potfolio and put all of the relevant terms into their resume anyway.
I think most aspiring data scientists would be better served going with more introductory texts and really understanding them. Maybe Blitzstein and Hwang's "Introduction to Probability" and then McElreath's "Statistical Rethinking" or Wasserman's "All of Statistics" for people who need more stats.
I'm not even sure what to recommend for developing good software judgment and habits. There doesn't seem to be a shortcut for that. Maybe "Fluent Python" or "Effective Python" for Python people? No idea for the R ecosystem.
I’d like to learn to do my due diligence, but without someone training me, it just takes so much time to learn things like git. I’d rather be recording more data and submitting my paper so I can get the hell out of here
I just assume anyone who calls themselves a data scientist is going to be a shit tier programmer who needs to improve over time at this point. The exceptions kind of prove the rule. Imposing test-driven discipline will cure some of the worst tendencies.
I don't have good references on stats and linear regression tier data science, but I'll take someone who understand the basics (I dunno, calculating useful moments from empirical distributions, feature selection in linreg) over some weenie who has some cribbed ipython file in his githubs who claims to understand Hastie.
I’ve gone out of my way over the years to make learning data science skills as approachable as possible for uninitiated (giving trainings, providing customized learning paths based on someone’s background, offering encouragement), and yet this is almost never reciprocated by engineer types. It’s always just, “data scientists can’t write production quality code”, with no explanation of what production quality entail, or without consideration of the fact that notebook-based data science can have advantages over perfectly modularized code with a battery of tests. See the comment above: “I'm not even sure what to recommend for developing good software judgment and habits.“. It’s like a chess coach admonishing their subject to simply “think harder”. Not helpful.
When curious and open-minded data scientists and software engineers work together, it can be magic. When people snipe at others for their “shitty” skills, it creates a petty and toxic environment.
This comment comes off as a bit of an admonition, but I would greatly appreciate a list like TFA for data scientists looking to shore up their fundamental CS and software development skills.
(PS — The first book I read when teaching myself R was R Inferno, so that ain’t it.)
Hey, it seems like you took this as gatekeeping or something. These skills can definitely be taught or self-learned, I've done it and seen it done many times.
My point was only that I don't know resources that can act as a shortcut (my actual word above), i.e. ways to skip over the longer path of gaining experience through long engagement with the topic. So maybe more like a chess coach saying they don't know any books that let a beginner jump ahead to being a more experienced player?
There are hundreds of past threads on HN about books to level up in software, so clearly some people have thoughts about this. I just don't know what to recommend a data scientist who needs these skills immediately.
Sometimes a rant on a topic brews in my head for weeks or months, and I will uncork it on a random passerby that brings up the subject—which happened to be you this time.
But, I’ve had coworkers who like clockwork sneer at anything a data scientist wrote. “Why did you do it that way?”. When asked for advice on how to improve it, they huffily say nevermind. It’s ingratiating as hell.
As an software engineer (2) is the worst part of working with data scientists. The amount of times a 'professional' data scientists want to launch non-code reviewed, non source controlled, works-on-my-notebook model or analysis is shocking.
https://www.amazon.com/Python-Data-Analysis-Wrangling-IPytho...