I ask because I do this kind of work as my J.O.B. and every skill he refers to I totally see the necessity of except this one.
I ask because I do this kind of work as my J.O.B. and every skill he refers to I totally see the necessity of except this one.
Representing features of a datapoint as a vector, pervades and populates every pore of this field. Without an understanding of linear algebra you wouldn't have support vector machines, no kernel methods, no neural networks, no perceptrons, no gradient descent methods, no Newton / Quasi-Newton methods, no multi-dimensional (or as they say in statistics, multivariate) Gaussian random variables, no matrix factorization, no Pagerank, no Markov chains, this list can go on and on.
Take the simplest of data science problems: you have one variable x and another variable y and you want to predict the value of y given x. Usually x is not a single scalar but n scalars (called a feature vector). Simplest thing you can do here is least squares and that is as linear algebraic as you can get. There many fancy ways of dealing with this problem but almost always it is reduced to solving a related linear system.
The bottom line is this: we understand very few things. Thankfully linear algebra is one of the few things that we do understand, so almost every analytical problem is reduced to this case (if, but locally) and then solved.
I would be very curious to know how you have been able to avoid linear algebra. It will give me a new and valuable perspective, because apart from "click button, didnt work? ok click the next button" data analysis I find it hard how one can do much data analysis without it. So please break my bubble, I will be thankful for it.
Canned packages often do not work out of the box. The knowledge of linear comes very handy in analyzing and debugging why is the model not working" "oh I see this matrix is near singular, thats why my estimates are off the park", or "oh these two variables are very correlated, that is why gradient descent is having so much trouble converging fast", "ah I see why I am getting NaN here" etc etc.
EDIT: darkxanthos, appreciate your comment. I would say it is a bit like driving. Knowing the internal mechanics is neither necessary nor sufficient, and hardly correlated with good driving skills when things are going well. But sometimes when things are not going as expected, it helps in debugging.
Let me try and pique your interest: Note that the decision boundary of naive Bayes is actually a linear function of the log conditional probabilities considered all independent, with LA you can now also consider the case that they have dependence. Consider updating multi-armed bandit problems, the updates are variants of gradient descent, and its nature is indeed characterized by the eigenvalues of Hessian of the thing you want to optimize. Consider K-means clustering, one way to get very close to its global optimum is to solve the same cost function using linear algebraic updates (called spectral graph partitioning). By trig I think you have the dot-product of two vectors in mind, the related analysis actually does not rely much on trigonometric properties but heavily on the linear algebraic properties, in fact this what allows one to escalate affairs from simple linear feature vectors to extremely non-linear ones because even though they are nonlinear in the data space in some other space they are linear so people do the math in that space (called the kernel trick although I find that term quite silly) ..This thing, linear algebra, lurks everywhere, I tell you :)
For example- Least squares regression. Totally use this. Even took a semester in college on just regression. The linear algebra underpinnings though haven't never been shown except for a quick blurb in my linear algebra text book. I still understand the concepts of fitting a model and when it's a bad fit (such as non-normal distribution of residuals, co-linearity) but the theoretical underpinnings are more fuzzy to me.
Representing features as vectors, sure. But that's also a pretty superficial use of linear algebra since from that point forward I'm using something on the trig side to compute results (at least in clustering).
I also tend to lean rather heavily on probability and bayesian approaches to many areas. So Naive Bayes classification is a love of mine, finding ideal parameter values given data coming in becomes an online updating multi-armed bandit problem to me (which also doesn't require explicit linear algebra). A lot of my work is also in experiment design and analysis and for this I use a mixture of Bayseian and frequentist statistical testing.
Canned packages out of the box with parameters to tweak that I can cross validate to evaluate how well my model is working. If I happen to venture out to other models I'm probably reading up on common pitfalls and how to test for them.
To me, it's entirely possible that the gap between you and I is due to experience and even just differences in training/learning (including but not limited to the quantity of it). These discussions are important for me since they help to inform my future learning aspirations.
In this scenario, only knowing about mechanics still wouldn't make you a great driver, but it would help you design a vehicle that works for you, and help identify vehicles that might blow up with you inside!
That having been said I use ML algos, and I suck at linear algebra. I often wish that my linear algebra classes hadn't been so dreadfully stale and boring.
The multivariate Gaussian distribution is a great example of this. It's probably the most fundamental and important distribution in statistics, and working with these distributions is pretty much pure linear algebra -- quadratic forms over vectors of parameters, eigendecompositions of covariance matrices etc.
Even for non-statistically-motivated data mining: any time you're optimising over a lot of parameters, it's likely that linear algebra (and, as before, multivariate calculus) will help. Linear algebra is as important to calculus over multiple variables, as plain old high-scool algebra is to plain old univariate calculus.
http://aix1.uottawa.ca/~jkhoury/app.htm
The linear algebra text by Anton (9th and 10th eds) has a huge section on applications of LA. Also this book does same for ODE's http://www.amazon.com/Topics-Mathematical-Modeling-K-Tung/dp...