Foundations of Data Science [pdf]
cs.cornell.edu
cs.cornell.edu
Paraphrasing this to data science: "Everybody wants to have software provide them insights from data, but no one wants to learn any math."
The top two comments here illustrate this perfectly. Anyone who is serious about learning data science will read this book and will not shy away from learning math. You can also learn about data pipelines, but that's not a substitute for what's in this book.
There are also a variety of other algorithmically focused machine learning books. They are also not a substitute for this book.
As far as I'm aware this book is one of a kind. This book covers ground that nothing else does, at least not in any comprehensive way.
The primary reason I dislike this term, though, is because all scientists are data scientists. Data is how the whole thing works.
Now that the term is diluted, I propose we use adopt the banking industry term "quant developer" to refer to that.
Back in the day, if you had small data you didn't need a data scientist. You needed a statistician. He'd do some shit in SAS/Python/etc, reading one CSV and writing another. Then the developers could run that on a beefy server with cron and push the CSV output someplace else.
At some point during the 2000's this stopped working - things became complicated and tangled enough that you couldn't just munge CSVs like this. You needed folks who understood the math well enough to come up with algorithms, and who also understood the computer science well enough to scale out. These folks were termed "data scientists". Banks call them "quant developers".
Very often you don't need a distributed system, or anything that isn't trivially parallelizable. Most of the time when I make money it's based on a CSV < 100GB - often it fits in ram. Nowadays I don't think it's misleading to call yourself a "data scientist" if you can only handle such data sets.
I recently learned that Facebook calls their business analysts data scientists. I guess it's just sexier.
There's natty bodybuilding for people who don't want to use PEDs.
THANK YOU for posting this comment. I absolutely LOVE these two quotes, and can't tell you how many times I could have used them over the years.
These two quotes are now a permanent part of my arsenal for responding to busy individuals who say they "want to understand" a new technology or idea, but what they actually want is for the new technology or idea to be explained to them using only simple concepts with which they are already familiar, so they don't have to learn anything new.
The problem is not math. The problem is the way math is explained. Most books do not go into the intuition and are very dry. People generally shy away due to the dryness of the material, not the content.
An N-dimensional vector space is not like R^3. A monad is not like a burrito, and >>= is not like chipotle forgetting to put chicken in your burrito, opening it up, adding chicken and wrapping in a new tortilla.
Math is a new thing rather than the same old thing with new syntax. It needs to be understood on it's own terms. If you are serious you'll learn it.
> In probability theory, the law of large numbers (LLN) is a theorem that describes the result of performing the same experiment a large number of times. According to the law, the average of the results obtained from a large number of trials should be close to the expected value, and will tend to become closer as more trials are performed.
This book: > If one generates random points in d-dimensional space using a Gaussian to generate coordinates, the distance between all pairs of points will be essentially the same when d is large. The reason is that the square of the distance between two points y and z ...
When anyone tells me that X is the ONLY way to do it, almost exclusively I have found them wrong - beyond data science. Formalism is a means of communication, not the end. You can always communicate ideas without formalism.
FYI, I am a Ph.D Computer Science and a practicing Data Scientist for many years now.
The problem is detailed in this wonderful essay by Paul Lockdhart
https://www.maa.org/external_archive/devlin/LockhartsLament....
Here's a simple concept: all polynomials with coefficients in a number system (field) have a solution (possibly in a larger field).
Do you know how hard it is to convey what exactly this means and what the consequences are? I like Lockhart's essay. It resonates with me. However, I don't see his ideas being useful in terms of learning advanced topics. At some point one has to get their hands dirty and slog through the material. Intuition will come with experience.
I never said it is easy.
> Intuition will come with experience.
And this experience that some people have already reached should be shared, not limited to a select coterie of people.
Except, I can hardly see that. At least, I see you trying - most people do not make an attempt at all.
FYI, I am a Computer Scientist.
90% of the time I use the java standard lib. Sometimes I need something almost the same, but a bit different. The difference between a middling developer and a good one is what happens in that 10% of the time.
The same applies to proofs. What happens when you need something that's almost like the theorem, but not exactly? Can you tweak the proof to apply to your case, or alternately recognize that it can't be fixed and you need to do something different?
Likewise, one can certainly apply a lot of mathematical techniques in a lot of situations, without proving new theorems. So there are edge cases... not the end of the world. Deal with it when it comes up.
The ability to determine when an approximation is good enough comes from a solid understanding of the underlying theory. For example, we understand that special relativity doesn't invalidate classical mechanics, because, in the limit, as speeds and energies become arbitrarily small, their predictions coincide. But, if all you have is a bunch of computational recipes, even the tiniest unexpected thing renders you unable to solve problems.
Anyway, I'm not saying that examples of applications are a bad thing. But proofs are indispensable. Even for “intuitive” people, proofs are necessary confirmation that their calculations will match their intuitions.
There is no proof of special relativity, we believe it because it makes experimentally verified predictions. Your example of limiting cases of special relativity is how I wish statistics texts were taught - it isn't based on the "proof" of special relativity. Of course I know when special relativity is applicable and when it isn't (when beta = v / c is close to one, special relativistic effects are important, and gamma = 1 / sqrt(1 - beta*beta) indicates how good the approximation is).
But I do agree that proofs are useful and needed, I think that complicated proofs could easily be placed in appendices and looked at after understanding why the theorem is relevant. Of course this is just my personal preference.
Right, that's because special relativity is a scientific theory. A scientific theory has two components: a mathematical theory (in which you can perform a priori physically meaningless calculations) and a physical interpretation (which turns the results of such calculations into predictions).
However, statistics isn't a scientific theory. It's just math, and like all math, it's about itself and nothing else. It doesn't make sense to “experimentally confirm a mathematical theory”, because, without a physical interpretation, math doesn't make any predictions about the real world.
> I think that complicated proofs could easily be placed in appendices and looked at after understanding why the theorem is relevant.
Then you're looking for books on applications. That's fine. But a book whose subject matter is a mathematical theory (it even has “foundations” in the title!) can't relegate proofs to appendices.
My problem with (most) maths texts isn't the proofs in and of themselves, it's just that they make so many assumptions about what you already know, and don't always state those assumptions clearly. And then too many maths texts, to me, fail to include enough expository text to explain the math and provide more of an intuition around what's going on.
That said, I also agree with you that more examples in terms of applications would be nice. Learning really abstract stuff in isolation is OK, but it's always nice to have a few (or more than a few) examples of how to apply the technique to something concrete.
Just skimming that part for the word "prerequisites" or "background" will usually find it.
Exactly. And a surprisingly large number of books aren't well-written, exactly because they don't include that bit.
For most well-studied topics, finding the books that are good enough is not that difficult.
Most actually-existing books simply write that the student should be "mathematically mature", by which they mean that the student should know all of multivariable and vector calculus, most of differential equations, much of optimization, and some amount of analysis. And also be able to do proofs, and also have good intuition.
In short, most "applied" math textbooks aim themselves at Master's degree students or beginning PhD students in math itself, to teach material which can actually be taught to anyone with a bachelor's in science or engineering (except possibly computer scientists, whose continuous maths degrade because we never use them).
Ironically, the first math textbook I've ever picked up that didn't assume far too much background was my real analysis textbook, because analysis teachers assume that they are teaching the gateway course to "real math" and have to educate little babbies who only just got done with calculus.
If you want to learn something difficult that's above your level, and learning that something includes learning a prerequisite subject that's also above your level, you simply start at the node in the DAG that's at your level and work through the prerequisites if you're serious.
But yes, there is no substitute for learning all the things you need.
One thing I am constantly frustrated by is the lack of mathematical rigor that pervade DS tools, and programming in general. Seems like everyone is happy to be given some programmatic tool and all the mistakes that where programmed into it. "You just need to deal with that." "You'll get used to the syntax." Yeah, but, I'd rather not, since it's often completely arbitrary and specific to a given tool. And, well, sometimes just doesn't make sense in the way something that has mathematical rigor built in might.
But, I often feel when I bring this issue up I'm shouting down an empty hallway.
I think that this kind of approach is going to get us pretty far -- see how far we got on building bridges with no more sophisticated insight in to gravity than "things fall down; some things are heavy" or material science than "some things hard and not bend" -- but I think that ultimately, our longer term goals like AGI will require that we build theoretical models capable of explaining our current engineering progress and predict a new, deeper valley of progress to move to.
Of course, that's a lot like saying making progress on physics paradigms requires explaining high temperature superconductors (or other currently open problems). It's actually pretty normal for a field to be a mix of engineers being ahead and theorists being ahead.
Mathematical rigor is really a sign that you're in established territory and explorers have moved on, and much of data science is still under heavy exploration.
edit: what i would love to see is a "category theory" for programming languages / paradigms such that moving between them is well defined. it boggles my mind that translating between two programming languages isn't trivial. if both are well defined, one should be able to translate one to the other precisely given one is willing to define the translations. there should be zero guess work or heuristics in the process.
How does machine translation do at poetry, in the sense of capturing the figurative meaning and stylistic elements, not just the literal meaning of the tokens?
I think there's a lot of room for mathematics education to be revamped to be more iterative and exploratory, which I think would make it appeal to a wider audience. I think that the mathematics community is actually beginning to pick up on this, with the publishing of python notebooks which contain increasingly sophisticated models as you read in to them and provide controls for readers to manipulate the parameter space of the models to explore the consequences of claims, theorems, etc. I think this kind of exploration of the numbers is a better way to learn a lot of ideas than the way I was taught in class, particularly things like statistics, where you can see distributions change as you manipulate parameters.
I figure that in a few years, as people my age who were mostly taught using older methods, but got a little exposure to explore-via-computer learning in university mature, we'll see an increasingly "internet-y" version of teaching methods appear, where there's a mix of video lectures, reading materials, discussion posts/blogs, and interactive models.
Computer programming and hobby crafts are among some of the first fields to make that transition, traditional academics are somewhere in the middle of the pack, and specialized topics like data science are keeping pace with the wider academic community.
tl;dr: I think your point is really just a temporary pedagogical fluke, as we take time to transition our teaching methods.
btw, I used to think like you before I got into math. Its a very valid albeit flawed viewpoint - to think that hey if only there was a javascript doodad somewhere so I could move a slider and manipulate the parameter space, I will instantly "feel" CLT in my blood instead of working through CLT as a dry formal math proof. The reality is, this sort of dumbing things down works maybe in 2D and best case 3D, but beyond that, its utility is rapidly diminished. Also, the results you get in the lower dimensions don't map into higher space. As yummyfajitas points out elsewhere, in a reasonably high dim space, all pairs of random points will be fairly equidistant, which is simply not the case for say 2D. So even the little bit of useful intuition you learn in the lower dims becomes quite useless as you scale up.
Proofs are best learnt by doing proofs. Math is not cinema.
The advantage of this approach is that it works. You can acquire solid intuition this way. But it is arduous as hell and also pretty one-sided - problems are limited to what an average undergraduate can do by hand.
Now, when I think about using computers in math education I don't think of some "javascript doodad" with moving sliders. I imagine instead implementing an algorithm or doing a numerical experiment, learning about possible pitfalls in the implementation, running the algorithm on some real data, seeing it fail etc. I think this experience is pretty valuable too and can nicely complement traditional approach.
As an example of what I am talking about, the book "Structure and Interpretation of Classical Mechanics" (a cousin to SICP) tries to teach classical mechanics with the help of computer programs. But this approach remains unconventional.
Now, I'm not a mathematician or a scientist of any kind, merely a dilettante textbook reader, but I'm not sure that approach is sufficient, or even most efficient to teach all subjects.
I think it arises naturally when you approach subjects where it's directly relevant: engineering, complexity theory, type theory, etc. and people who will need those insights will get them in due time.
But yeah, broadening the audience could help people get a better appreciation of the connections between pure and applied math.
That's really kind of pathetic, and a disservice to your view.
I'd be happy to speak with you, though, if you want to actually read my comment and address a (charitable) interpretation of what I said, rather than your strawman one.
For many industries, this problem is not well-solved at all. And it's not trivial. It's one thing to have great models, it's another to have a good way of making people smarter with it.
However, having said that, I still consider myself a very productive DS, and I get stuff done. I'm not going to ever be on a research team at Google, but those aren't the only kinds of DS jobs out there.
And I'm not just plotting stuff.
Mastering basic theory, on the other hand, needs a coherent and structured study-plan which requires extended focus and single-minded emphasis (at least for most folks).
Logically yes. But it's amusing how many scienty types (PhDs et al) who know all the fundamentals in theory can't do such practical tasks if their life was depended on it.
It's like they thought the theorems and abstract objects they've learned would never be encountered in the wild.
There's probably a scale at which that works better, two pizza data science teams for example.
So I can imagine a pipeline developed for use with well established social media APIs or standard scientific experiments that have been in used for decades. But it is hard to imagine a pipeline that can handle amorphous emerging high throughput instruments/methods.
It seems like a natural evolution of DB admin for very large scale noSQL and DBs, esp. those with unstructured data and often non-commodity architecture. I've seen numerous companies looking for such folks and I suspect demand will rise.
IMO, it's not a job you'd want to outsource. The role is too mission critical, the skills not predictable enough to be a commodity, and the penalty for screwing up is too great.
Also, chapter three about SVD, is in https://jeremykun.com/2016/04/18/singular-value-decompositio...
and https://jeremykun.com/2016/05/16/singular-value-decompositio...
the advantage is that you have the python code available.
https://jeremykun.com/2015/04/06/markov-chain-monte-carlo-wi...
The book seems to be interesting.
For example: - What is the distribution of the residuals, how does it change over time as data comes in. How Gaussian they are(or not), analyzing weird/oddities especially around the tails - What kind of features offer the most significant signal to the model and which ones are not.
These skills are even applicable to SVM and other classification analysis.
So just being honest, the Appendix is still rather terse and advanced for me. Does anyone have suggestions for prerequisite readings that would help getting someone prepared for this text?
- Calculus covering Derivation, Integration, Multi-Variate
- Linear Algebra
- Differential Equations (May not be relevant here)
'Discrete math' is also useful.
Some of the derivations in the book you can take on faith and not fully prove out to save time, but you should feel 100% comfortable/confident with the notation used.
Let me know what's confusing you and we can try to figure out what you are missing.
Ok I'm being extremely facetious here but it still shocks me when people comment in machine learning or data science threads asking about how they'd go about picking this up but you can clearly tell they have no formal background in the sciences. I guess the only reason it angers me is because if they had just done a CS degree instead of trying to "hack it" none of this would seem like magic.
1. Book of Proof by Hammack (http://www.people.vcu.edu/~rhammack/BookOfProof/)
2. Calculus by Spivak
3. Linear Algebra Done Right by Axler
Be prepared to work through all (or at the very least only the odd numbered) exercises. If you can't stomach that or find that life gets in the way of you completing even these very basic books, you do not have the time or discipline required to advance in mathematics.
there are now "warm up" books (alcock) as well as even more basic real analysis books (abbott, about half the length of spivak).
I assume multivariate calculus and linear algebra?
Notably missing are causal inference, experiment design, and many topics in statistics--causal inference being one of the primary things we'd want to do with data.
Some context: I did my undergraduate degree in Economics (in a pretty math intensive university), have been working in marketing for the last 2 years and want to go back to do work in something more analysis centered.
It has a lot of topics on applied math.
Some of its main topics are linear algebra, probability theory, and Markov processes.
Really, the book just touches on such topics. Usually in college each of those topics is worth a course of a semester or more. So, what the book has on such topics is much less than such a course. E.g., for linear algebra, the book gets quickly to the singular value decomposition but leaves its treatment of eigenvalues for an appendix and otherwise leaves out about 80% of a one semester course on linear algebra. Similarly for probability and Markov processes.
Some of the topics the book has or touches on are unusual with, likely, few other sources in book form. E.g., early on the book has Gaussian distributions on finite dimensional vector spaces where the dimension is larger than is common.
So, for the topics rarely covered in book form, the book could be a good reference.
For topics such as from linear algebra, a reader might get misled without an actual course in linear algebra from any of the long popular books, e.g., Halmos, Strang, Hoffman and Kunze, Nering or more advanced books by Horn, Bellman, or others.
Usually in universities, probability and Markov processes quickly get into graduate material with a prerequisite in measure theory and, hopefully, some on functional analysis, e.g., to discuss some important cases of convergence.
So, the book seems to have some good points and some less good ones. A good point is that the book is a source of a start on some topics rarely in book form. A less good point is that the book gives very brief coverage of topics otherwise usually covered in full courses from popular texts.
A student with a good math background could use the book as a reference and maybe at times get some value from the coverage of some of the topics rarely covered elsewhere. But I would suspect that students without courses in linear algebra, probability, etc. would need more background in math to find the book very useful.
E.g., early in my career, I jumped into various applied math topics using very brief treatments. Later when I did careful study of good texts with relatively full coverage, I discovered that the brief treatments had been misleading. E.g., no one would try to learn heart surgery in a weekend and then try to apply it to a real person. Well, for applied math, maybe learning singular value decomposition, etc. in a weekend might not be enough to make a serious application.
It is good to see a book on applied math try to be a little closer to real, recent applications than has been traditional in applied math texts. I'm not sure that the being closer is crucial or even very useful for making real applications, but maybe it will help.
* High Dimensional Space
* Best Fit Subspace & SVD
* Random Walks & Markov Chains
* Machine Learning
* Massive Data: Streaming, Sketching, Sampling
* Clustering
* Topic Models, Hidden Markov Process, Graphical Models, and Belief Propagation
So yes, this book that is avaialble for download from a TURING Award winner's web page IS a Computer Science book for Data Science.
Edit: I'm not sure why email is not showing up. Here is an old site with contact info: https://sites.google.com/site/thomasroderick/
"Data Scientist" is for Ph.D.s or sat least Masters, those are the less common jobs.
At least that's the theory, we all know what happens with titles... anyway, the point is that we can safely assume that this particular article "Foundations of Data Science" refer to the actual "scientist" role.
https://www.datacamp.com/community/tutorials/data-science-in...
https://bigdatauniversity.com/blog/data-scientist-vs-data-en...