Principal Component Analysis for Dummies
georgemdallas.wordpress.com
georgemdallas.wordpress.com
In case you're not familiar with them, the basic idea is treating a image of a face as a very high dimensional vector, and then doing what amounts to PCA on a collection of them. I'm leaving off a few steps, but the resulting eigenvectors converted back into images helped me grasp what was going on in a much more intuitive fashion.
This is also closer to it's actual implementation: while it's true that you do technically need the eigenbasis of the covariance matrix, you should not actually form the covariance matrix to get there...
you should not actually form the covariance matrix to get there...
In case anyone's wondering why: it's not only because it takes extra time. The main reason is that computing X*X^T can bring numerical instability, where a direct SVD(X) would work just fine.Also
(X'X)g = lg
(XX')Xg = lXg
So with n>>p and n large, you may be able to fit X'X (p^2 entries) but not XX' (n^2 entries) in memory (and get the same eigenvalues and related eigenvectors).
That's not how I'd explain it. The 'sought result' and its reflection are not two different things. They are the same thing, differing only in a trivial detail of orientation. The 'negative' of the result conveys exactly the same information as the result.
EDIT: By the way, you can make it readable with:
document.getElementById("wrapper").style.backgroundColor='white'
FTIR data analysis is a fantastic example for PCA analysis -- each principle factor ends up (probably) being the spectrum of one of the major real physical components. But this is maybe too abstract?
A less abstract one might be a distribution of test scores. Your actual dataset is "number" versus "score", and you could show two gaussians, one at a low number and one at a high number. Then you could show that across three exams, you always see the same scores, but with different intensities. That would let you compute that the principle components are those two gaussians. Then you can hypothesize that each group is a collection of students that study together, and so they get similar scores. Or something like that.
Anyway, no intent to be a wet blanket. It's a nice writeup, and it is nice of you to share.
[1] http://www.cs.otago.ac.nz/cosc453/student_tutorials/principa...
You can shortcut the whole process by finding the smallest non zero eigenvalue/eigenvector pairs of the graph laplacian (Fiedler vectors). You need to use a sparse solver that can find the smallest values/vectors instead of the larges (like LOBCPG) but that is faster anyways.
PCA is a form of (or at least related to) correlation. With standardization the resulting transformation hihlights variables in the original data that are most highly correlated. Without standardization you're visualizing covariation. Unlike correlation, covariation is influenced by the magnitude of the variables.
By standardizing, you control for differences in the magnitude of the variables, and focus on their inherent variation instead.
However, let's set that aside. I apologize for being a bit obfuscatory. My point is: If this is the case, then the explanation in the OP is totally misleading, because your data shouldn't look like an ellipsoid, but rather a circle. PCA should only be used in situations where there is a reason to believe there is a mechanistically justifiable "hidden value" that underlies otherwise uncontrolled "independent variables", thus making a dimensional reduction reasonable.
This is not at all the situation that the OP goes over in the first part of the post.
This is easily grasped with a 2d example, despite the fact that PCA makes no sense with only two variables.
http://www.amazon.com/review/R16RJ2PT63DZ3Q/ref=cm_cr_rev_de...
It seems like PCA would already be a method that would only mean something if it was applied to comparable dimensions. What would transformed, dimensioned variables mean anyway? Chart A=mass - 3charge by B = mass + 2charge. What could a correlation mean.