Linear compression in Python: PCA vs unsupervised feature selection
efavdb.com
efavdb.com
> The printed lines above show that both algorithms capture more than 50% of the variance exhibited in the data using only 4 of the 50 stocks.
Based on the sklearn PCA documentation [1] this has nothing to do with the coefficients on individual stocks, and for PCA should read more like: "[...] capture more than 50% of the variance exhibited in the data using only 4 components [...]" which is not the same thing.
1. http://scikit-learn.org/stable/modules/generated/sklearn.dec...
This kind of interpretation kind of falls out of the math (the eigendecomposition/SVD/covariance matrix interpretation of PCA in particular).
(Ed: yeah, that's just a sample of the book but has a large bibliography at the end.)
>>> selector.ordered_cods
[0.43298218, ... , 0.5068577, 0.5068577]
Would you think this a problem/bug?If interested, you can find some detailed examples in a tutorial below https://github.com/EFavDB/linselect_demos
The title seems like it has the form "<Specific method> vs <Broader category method fits in>".
Not disagreeing with you, just spitballing.