A model overfits if it is unnecessarily COMPLEX for the training data.
If there is bias in the training (and validation and test) data that allows a SIMPLE model to fit the data because of a spurious correlation, that is not overfitting.
A model overfits if it is unnecessarily COMPLEX for the training data.
If there is bias in the training (and validation and test) data that allows a SIMPLE model to fit the data because of a spurious correlation, that is not overfitting.
More specifically, politically undesirable correlation - as in, "it's there, but its existence upsets some people". It's pretty obvious and self-evident that there are meaningful biological differences related to age, sex, and other demographics. Whether or not they're clinically relevant for a specific diagnosis under question is one thing, but they are clinically relevant for great many diagnoses; trying to "de-bias" reality here will only lead to unnecessary suffering and loss of life.
Now you can say that this is perfectly fine and represents the most likely real-world use case. Or you might prefer a model that looks at the image only, with the implicit assumption that this "forbidden knowledge" will be added by human doctors later on in the pipeline. This is beneficial because the "forbidden knowledge", such as whether patients from Hospital A always have bone cancer, might change overnight! Imagine the hospital gets assigned a new name in the system and the prediction is shit now.
This second, "unbiased" AI system will always have a worse performance, because you lobotomize it when you kill the forbidden knowledge with a sledgehammer. This study just showed that "group fairness" is at odds with optimal predictions" and how much it is at odds.
PS: You might even prefer a society where everyone is worse off, but every protected group is equally bad off. You'd also ban the humans from applying the forbidden knowledge. Whether that is desirable, is, of course, out of the scope of the paper.
These models are only looking at the images. They are inferring demographics.
Your knowledge guides, but it also doesn't (or shouldn't) blind you.