https://www.jmlr.org/papers/volume3/perkins03a/perkins03a.pd...
and gain-based selection (using the improvement of the objective), see the appendix of:
https://aclanthology.org/J96-1002.pdf
We used grafting for parser feature selection, for which it worked quite well:
Another commenter mentioned L1 regularization, which is useful for linear regression. You wouldn't use it for all classes of problems. L1 regularization has to do with minimizing error of absolute values, instead of squared errors or similar.
I skimmed this article and thinks it's accessible: https://www.kdnuggets.com/2021/06/feature-selection-overview...
PCA is a form of dimensionality reduction, but it doesn't select features for you.
Not quite. It has to do with minimizing the sum of absolute values of the coefficients, not the error. The squared error is still the "fidelity" term in the cost function.
Lasso is also known as L1 regularisation, and it tends to set the coefficients to a bunch of features to zero, hence performing feature selection.
Note that if two predictors are very correlated, lasso may pick one mostly at random. Obviously one should do CV and bootstrapping to ensure that the results are relatively stable.
In general though, there's no real substitute for domain expertise when it comes to selecting good features.
edit: lasso is L1, not L2
besides techniques mentioned in other posts (l1 regularization) sequential feature selection (backward and forward) is quite common
This is fine if you have test/validation sets, but never, ever report p-values on the result of such a selection process, as they are incredibly biased.
This one in particular compares a few methods on 38 datasets and has some Python code: https://blog.kxy.ai/adding-feature-selection-to-any-model-in....
Disclaimer: I wrote the original blog post.