What is more, Leo Breimain wrote on his website: "Random forests does not overfit. You can run as many trees as you want" https://www.stat.berkeley.edu/~breiman/RandomForests/cc_home...
Deep trees will fortunately overfit your dataset.
Any binary tree of depth log2(P) can completely separate your P points.
False claims as they maybe, these are claims I've seen in at least two of the most commonly studied statistical learning text books, so given that it makes sense and that it's in the text books, it seems reasonably not false to me. Someone else posted that if too many features or data points are very similar then it will overfit, and that totally makes sense. Whatever you say doesnt. Clarification would be useful.
I have an explanation here why reducing variance is not the same as reducing overfitting: https://news.ycombinator.com/item?id=20089890
Boosting is sequential, and relies on early stopping to control the magnitude of bias.
Sounds like overfitting.
Example, you could have a tree that individually segments each data point and memorize your dataset. That's the definition of overfitting.
The « variance » won’t just magically vanish as you average things out[1], you need to change the scale and check out the asymptotic law of your estimator (CLT, Kolmogorov-Smirnov… etc.) and confront it to your data.
[1] the variance of the estimator itself vanishes thanks to LLN (in case of convergence), but that’s not actually the quantity of interest
Edit: don't get me wrong, I'm not saying that RFs are good or bad, just reacting to the bias/variance thing.
Think of the bias/variance tradeoff as a spotlight, and we are shining the spotlight on a bunch of cats, who reflect back the spotlight when their eyes are open. Eyes are open or closed randomly. Cat eyes are either green or brown. We want to know the distribution of cat eyes in parts of the population, which in general is an even 50/50 split. We determine the distribution in a certain location by taking the average of the eyes we see.
If variance is large, then the spotlight is very large, and we don't learn anything because we just average the entire population.
If the spotlight is small, then we can learn something, but only if there are enough samples in the region we shine the light.
So, what if we start with a large spotlight, and then when we see a region with a large number of open eyes of one color, we narrow the light down to just that region? Won't that allow us to avoid overfitting, while maximizing our ability to learn?
It unfortunately does not, because with a large enough population that is evenly distributed, there will always be pockets that exhibit what appear to be a pattern, but is just an accident of which cats happened to open their eyes.
This scenario of starting with the spotlight large and then zooming into a patterned region is the same as reducing variance with the training data. With a large enough dataset it is always possible to find these accidental patterns and then zoom into them by reducing variance.
https://stats.stackexchange.com/questions/20714/does-ensembl...
In my line of research I am frequently trying to use high dimensional data, but with few examples (<100 per class). Thus methods like SVM are used. I've been thinking about how I might leverage my sample to artificially simulate new training examples via pairwise warping of images within each class, with the assumption that informative features will be preserved with warping.The training examples within class are already quite variable, so I don't think a little increase in redundancy will hurt me much..but I am not sure.
Without knowing more concretely, do you have thoughts on such a strategy?
Data are 3D brain images and classes are disorder groups.
You can also try more generic upsampling techniques, like SMOTE, which should be easy from python or R. It's never actually helped me, but I assume it's useful somewhere.
I suspect at some point you're going to need to take an axe to some of your inputs, preferably based on human priors rather than a sketchy feature-selection process.
SVM's are great, but once you get past linear boundaries there's enough tuning complexity that I'd rather use that effort tuning a GBM. That's largely because of tooling though; I know there are modern SVM libs, but I haven't used them. Definitely try a random forest if you haven't!