A First Course in Linear Algebra
linear.ups.edu
linear.ups.edu
IMO this is the wrong way to understand linear algebra, and the typical path of a lot of introductory text books. A better text: Axler's "Linear Algebra Done Right"
PD: If you really do like or need the algebraic approach, Shilov's book is the way to go after this course... I think the first chapter is about determinants and he builds the rest of the book from there. It also includes a chapter on tensor algebra.
As a math professor, I think you can get by without defining vector spaces -- I think it is a little bit difficult to come up with examples, other than R^n or C^n, which seem motivated and interesting to the beginner. The truly important example (IMHO) is R^n without a choice of basis -- but I think this can only be well motivated after you've seen a lot of linear algebra, not before.
Mechanics of matrix operations are not pleasant to teach, they make the subject seem like a bunch of contrived and confusing examples. Same for the "row echelon" stuff. Alright, you now have an algorithm for solving systems of equations... but presenting linear algebra as an algorithm sells it short.
What is really important in my view is the geometry of the subject, which is already very interesting in two dimensions. Problem: Here is a linear transformation, given as a 2x2 matrix. Draw a picture which illustrates what this does to the plane. If you ask me, this is much more important than most of the crap that gets taught and tested in most courses in linear algebra.
My introductory linear algebra class also used polynomials over R, which led to examples like using projection to construct a polynomial approximation of a non-polynomial function.
Now that I'm out of school, I really want to learn how to properly use matrix algebra. I don't care about HOW to find a determinant or HOW LU decomposition works as I've already done those, but rather why I would want to perform that operation on the matrix in the first place.
The ability for matrices to model real processes is something that really fascinates me.
Do you have any other good reading material for someone like myself?
I'm afraid I don't personally (I'm into abstract and theoretical math, which I'm guessing is not your cup of tea). But I should dig something up before I teach the subject again, so I would be as interested to read replies as you.
Gilbert Strang is an applied mathematician and also the author of a popular linear algebra book; I would guess that his books might interest you. But this is speculation, I haven't read any of them myself.
So what is a good Linear Algebra textbook from that perspective? Assume someone has completed a course in Axiomatic Set Theory and is conversant with proofs etc, but now wants to get into Linear Algebra. Any textbook suggestions?
I would suggest trying these videos: http://www.stanford.edu/~boyd/ee263/videos.html. The prerequisites are very low and a main focus is on interpreting the abstract concepts in applications.
Another resource I've used is Patrick JMT's excellent videos found at: http://patrickjmt.com. He really goes over problems slowly. Best math teacher I've ever had.
and/or
http://people.math.gatech.edu/~cain/textbooks/onlinebooks.ht...
http://www.reddit.com/r/newschool/comments/wmu5q/read_this_h...
And it's something you slowly forget if you don't put it to use.
What do you consider to be the high point?
Bravo.
btw, I couldn't use Singular Value Decomposition in numeric.js for PCA because the method, numeric.svd, uses the "thin" algorithm, and throws an error if there are more columns than rows. I calculate way more features (50-200+ columns) than I have training samples (rows, 10-30 written manually). without svd I had to use the "covariance method", which I guess can sometimes present approximation issues, but seems to be working well for me.
The purpose of the PCA is dimensionality reduction (google "curse of dimensionality"). I had used Mahalanobis Distance as a p-value score to detect outliers (p-val < 0.05), and it worked well when there were only 6 features. Curse of dimensionality makes MD useless when there are 50 or 100 features, and PCA reduces them to 3-10 features which carry the most information with the rest approaching zero. So I do MD on the projected (reduced) features, and its working great.
If I do it all over again I might try a "one-class SVM" (which, sadly, I had not heard of until only recently and very late in the project). SVM's are non-linear like most machine learning algos, but the linear PCA can still be used to complement other methods, eg to do pre-processing before feeding to a neural network.[1]
For deeper linear methods, check out the PCA-related extensions like Fisher Linear Discriminants or Projection Pursuit.
1. "Many neural networks are susceptible to the Curse of Dimensionality though less so than the statistical techniques. The neural networks attempt to fit a surface over the data and there must be sufficient data density to discern the surface. Most neural networks automatically reduce the input features to focus on the key attributes. But nevertheless, they still benefit from feature selection or lower dimensionality data projections."
1. Faloutsos et al, Quantifiable Data Mining Using Principal Component Analysis, 1997
I agree, it's very useful. He gives nice explanations of some of the decompositions and why they're useful.
It is very good and explains all the proofs and definitions very well.