“‘Principal Component Analysis’ is a dimensionally invalid method”
amazon.com
amazon.com
If PCA is not being used to reduce the dimensionality of multivariate data, this might be invalid. There are other uses of PCA besides working on data (that image reduction technique using SVD comes to mind) that he might be addressing.
If you want a treatment of PCA in a respected text (at least by the statistical community--not sure what the ML people think ;) ), look no further than Hastie, Tibshirani and Friedman's Elements of Stastical Learning.
http://www.stanford.edu/~hastie/local.ftp/Springer/OLD//ESLI...
Is this similar to "whitening"?
For any given problem, you should generally only use methods that obey the symmetries that your data has. This doesn't mean that PCA is invalid as a method, it's just only valid on a data where the scaling symmetry does not apply. Another example would be fitting a line through the origin for temperature data. That's invalid, because the result you get depends on where you define your zero (but as before, it might be valid if that zero has special significance in your context). Does that mean that fitting a line through the origin is invalid for any data set? No.
In other words, the same criticism could be applied to any given method. Just choose any symmetry that the method does not respect, and then declare it completely invalid. Hence we cannot dismiss a method outright purely based on this reasoning. For example PCA is perfectly valid on unitless data. What's even stranger is that the author does like neural neworks, which are certainly not dimensionally valid, heck they probably don't satisfy any real world symmetries. This is also a case where it can be OK to use a dimensionally inconsistent method. As long as it works, it works.
So what ? The same is true of most machine learning methods including neural networks and SVM's, you just have to use the same units consistently.
I don't see the problem with PCA if it is used sensitively to identify the approximate dimensionality of "pancakes" embedded in a larger space.
Neural networks are affinity invariant. You can rotate, skew, translate the input data however you want, the optimum stays the same. Same for SVMs.
Edit: Ah right - http://en.wikipedia.org/wiki/Whitening_transformation - let covariance matrix be I.
PCA tries to project to the subspace that preserves as much distance in the input space possible. If you multiply a coordinate in the input space by a factor of 2, it will contribute relatively more to the distances, and hence change the fitted projection beyond just a scaling factor.
Edit: Ah right - http://en.wikipedia.org/wiki/Whitening_transformation
The problem with PCA is that everyone knows about it and kind of understands it, and thinks that it'll magically tell them something interesting about their data. Especially when you can take the first three components and make cool-looking 3D plots...
By the way that's worded, it sounds like a case of sensitivity rather than the method. If you change one of the variables to a unit that is way out of scale, then it's quite possible that the results of PCA, and many other methods, will change. But that's because it's not scale invariant, and so if you want good results you need to present your variables in the same units, and/or in some normalized format (zscored, etc.) where the scale of one unit doesn't blow the others out of the water.
These methods are not magical, and they are not intelligent. They do not know what they're looking at, so it's your job to feed them something reasonable.
That said, if you take some physically grounded data and change units, you won't recover different eigenvectors as long as your input data is good. The physics doesn't care about the units, or the coordinate system, or whether you use python or anything else.
I don't think that's accurate. Consider a set of points distributed along a line in 2D. If you do PCA on these points you will find that the 1st eigenvector points along the line and the second is orthogonal to it.
If you now rescale the axes so that x is measured in meters and y in light years the slope of the line will change and so will the the 2 PCA eigenvectors.
However the relationship between the 2 eigenvectors and the distribution of the data points will remain the same. The first eigenvector will still point along the line and the second will still be orthogonal to it.
In machine learning one is interested in the distribution of the data not in whatever units they happen to be measured in, hence I don't understand MacKay's objection.
Principal component analysis is predicated on a choice of inner product, since component directions are always chosen to be orthogonal. It's not clear what orthogonality could mean in a 2D plane where one direction is measured in inches and the other in tons, so naive PCA isn't appropriate in such a case.
Others have mentioned in this thread that "whitening" the data before PCA fixes this problem, by removing cross-correlations. Presumably, in that case, the notion of orthogonality is taken from the statistical properties of the data. (Maybe it normalizes physical units like inches to the standard deviation of the data's distribution in inches?)
But this scaling also resolves the issue for PCA. So, I don't see much difference between autoencoders and PCA with regards to original post's "dimensional invalidity" concern.
If anything, the scaling options you mention suggest "dimensional invalidity" isn't a big deal in practice for either method.
1. Reduce overfitting
2. Train models faster
3. Take advantage of unsupervised data
Regularization handles case (1) quite well. It can also be used in conjunction with most methods of dimensionality reduction such as PCA/auto-encoders.
PCA covers all three, but isn't as effective at dimensionality reduction compared to auto-encoders.
Auto-encoders tend to yield a better compression than PCA but take more time to train and produce output that's harder to understand. There is a bit of analogy here, auto-encoders are to PCA what neural nets are to linear regression.
[0] Feature selection, L1 vs. L2 regularization, and rotational invariance, Andrew Y. Ng. In Proceedings of the Twenty-first International Conference on Machine Learning, 2004. http://ai.stanford.edu/~ang/papers/icml04-l1l2.pdf
[1]"Augmented Implicitly Restarted Lanczos Bidiagonalization Methods", J. Baglama and L. Reichel, SIAM J. Sci. Comput. 2005.
The Hotelling transform is interesting in that it achieves optimal energy compaction, but has little practical value since it needs to be constructed anew for each dataset (usually, images).
Just to pull your thoughts away from massive, unfiltered big data :)
Considering that PCA works fine in the case that the dimensions form a proper vector space (e.g. geometrical space, stock market returns, temperature, etc, etc ,etc), it seems questionable to completely dismiss such a useful and historically important method.
load hald;
coeff = pca(zscore(ingredients))
ingredients(:,1) = ingredients(:,1) * 2;
coeff = pca(zscore(ingredients))
Magically, you get the same result regardless of if you change the units of one variable.Like all problems in statistics, it ought to depend on the specific task at hand. If there is some a priori reason to use the original scale (or a different re-weighting), it ought to be used. In general, PCA on correlation matrices is much preferred for exactly the reason you mention.
snark.
If you claimed to accept the null in an intro statistics class, you'd probably be failed.
If you want to argue for why "accept" is materially different from "fail to reject", feel free to do so - but I suggest that the chasm is by no means wide.
Apply PCA after non dimensionalizing any system. Read More here : http://en.wikipedia.org/wiki/Nondimensionalization
In other words whitening the data before applying PCA should result in the same eigenvectors expressed in the original coordinate system.
On the other hand if you're talking about the eigenvalues for the whitened data they're all 1.
So I'm still not seeing what whitening adds to PCA.
Just rescaling each dimension of the original space so that all dimensions have unit variance, without doing any rotations, may change things, but I don't think that's what is usually called whitening (according to Wikipedia).