Foundations of Data Science by John Hopcroft [pdf]
cs.cornell.edu
cs.cornell.edu
Why exactly?
I've found that being able to walk through a solution can be quite illuminating. There's also an exploration/efficiency tradeoff. At some point you can't spend any more time thinking through a problem (because life) and being able to work through a solution still brings many (if not maximal) benefits.
That said, most likely it's just to enable homework grading.
"There is an enormous increase in content when solutions are included. I trust my readers to decide which barriers they will attempt to leap over and which obstacles they will walk around. This often invites the objection that I am spoon-feeding my readers. My reply is that I would love to be spoon-fed class field theory, if only it were possible. Abstract mathematics is difficult enough without introducing gratuitous roadblocks."
Solutions are one way to solve a problem. Providing an answer dissuades students finding novel answers for themselves.
here they illustrate a nice connection between large deviations and convex geometry - and have beautiful pictures.
so what are examples of high dimensional vector spaces? the set of all your customers (hopefully in the thousands!!) and the vectors include all of their transactions. there are many other examples
I'm biased, but I think that data science is statistics, and therefore the foundations of data science is statistics and probability theory. If you want to understand data science at a fundamental level, I would suggest taking courses in these areas.
The appendix covers Probability and Linear Algebra.
Statistics is a powerful lens through which to view all data science. E.g. supervised learning is building a model of the conditional probability P(y|x). Again, I am biased, but I think that methods that do not have some statistical interpretation are unlikely to be useful. E.g. if we take the graph of Facebook users and apply some matrix decomposition algorithm, who cares? What can we do with this decomposition? What does it predict?
E.g. people used to say "neural networks are a simple, flexible functional form for y = f(X,theta)". This turned out to be wrong: SGD training of neural networks has more advantages than the flexibility of the functional form. But it was a good hypothesis and starting point.
SVMs and decision trees have no statistical justification I know of. Low rank matrix approximation and k-means are justified by latent variables and non-parametric kernel methods respectively. I agree these justifications came after the fact, but they do give a way to understand how these models work.
Most importantly, all of the small tasks surrounding training a model are purely statistical, e.g. cross validation, different measures of accuracy, handling endogenous variables, etc.
Are you trying to argue that statistics is its own field and not built upon mathematics?
If I have not studied statistics I will not have a solid foundation in statistics, even if I have studied probability and linear algebra.
The only books that have ever felt coherent to me start with p(data, unknown) as being an approximate model of some domain. Everything then follows smoothly as inference, modeling, and computational methods or shortcuts.
In the end though the only thing that matters is whether a particular procedure has predictive value. Even with the most principled Bayesian analysis the model or prior are strictly speaking wrong (usually both), because the real world is complicated, so even there you're only left with testing predictive performance in practice. Careful probabilistic modeling is only valuable insofar as it lets us focus our search for inference procedures on those that are more likely to work in practice. A good counterexample is deep learning, which does not (or did not) have a solid probabilistic justification, but works extremely well in practice.