Foundations of Data Science (2018) [pdf]
cs.cornell.edu
cs.cornell.edu
Here are some actual thoughts about the book:
- Having actually read a few chapters of this book, I find it to be poorly written on the whole. The chapters are extremely verbose and written in a bit of a circuitous and winding style. I find the explanations to be strained and not very pedagogical. They might be insightful if you already understand the material well, but since this aims at being a "foundations" book, the style of writing seems to miss the mark. I think these guys might have been the wrong people to write this book.
- I don't think the selection of topics makes any sense. It's a bit of a weird mish mash of topics which are arguably more or less fundamental, and have been popular at different times in recent years in the applied math, statistics, computational science, machine learning, and electrical engineering communities. If you are a student, it seems unlikely that you will benefit from taking a course taught out of this book, since you would be better served by taking more focused courses on the individual topics from this book that happen to be relevant to your research. If you are a researcher and need to learn one of the topics in this book, it's unlikely that this book is a better reference than any of the many existing books on these topics (or review papers). If you are self-studying, I think you are bound to fail with this book. You will benefit from something which is more thought through pedagogically.
Ultimately, I don't really see the point of this book.
Most of the topics covered are interesting in their own way, but the pieces just do not add up -- it has the flavor of "every author contributes a chapter, ordered in round-robin style".
In particular: opening up with a whole chapter on the odd properties of high-dimensional spaces ("all the mass is at the edge", etc.) is interesting, but it doesn't work pedagogically. I've heard various talks, over the years, that present this family of results -- and it's nice to see so many of them collected here. But not as an introduction.
I'm doubtful that the subject of high-dimensional geometry (interesting though it is, our intuition in #dims = 3 does not scale correctly) provides frequent insight into why algorithms work or fail. (Happy to hear notable counterexamples on this.)
Name your favorite ML surprise: that weak learners can be converted into strong learners by boosting, that ANNs with 10^6 parameters can be trained with 10^4 data, that auto-encoder architectures work so well. If Chapter 1 is trying to shock the reader out of their complacency, that's probably where to start.
Also, this sentence from the last paragraph of the Introduction does not inspire confidence: "The term “almost surely” means with probability tending to one." Oops, that would be "equal to one."
The one example I know of re: insight is L1 regularization. The "spikiness" of the L1 unit ball in R^n for large n can be thought of as encouraging sparsity when using the L1 norm to regularize optimization problems.
"Foundations" in the mathematical jargon doesn't mean "introduction to". It's an illustration of several foundational topic in data science, which is markedly different from an introduction to the field.
E.g. "foundations of computer science" is not "CS 101".
In this sense (and in this sense only) I don't think it misses the mark.
I disagree on the verbosity, statements have full proofs but that's expected from a "foundations" book. CLRS is waaaay more verbose, for instance.
I agree that I don't see a strong connection or logical path in the selection of topics.
Edit: Also agree about the comments. Wow this thread is bad. The "fuck math" crowd must have awakened.
Sounds like an overview. I don't think any "foundations" book is intended to be the one and only source of information. I think it's supposed to give you an idea of where you want to go next.
If you're in the research community, you won't need to be told where to go next. You will be surrounded by people who will know, and you will pick it up by osmosis. If you don't, you won't last long.
If you're outside the research community, knowing where to go next is unlikely to be of much use.
If you work at a company as an engineer, it's unlikely any of the material in this book will be of much use to you.
How does a Data Engineering book cover topics on Data Science?
Edit: In my view, dificult to gain much from this.
https://www.frontiersin.org/articles/10.3389/feart.2022.1011...
but I guess the cool kids are using something new.