Study reveals why AI models that analyze medical images can be biased
medicalxpress.com
medicalxpress.com
So the training data is probably hugely biased, and the models will learn to predict the training labels as opposed to any magically correct “ground truth”. And internally detecting demographics and producing an output biased by demographics may well result in a better match to the training data than a perfect, unbiased output would be.
Its much easier for physicians to blame technology than their profession.
My understanding is that doctors may unconsciously do this as well, ignoring a possible diagnosis because they don't expect a patient of a certain demographic to have a particular issue.
I would expect radiologists who practice in very different demographic environments would not do as well when evaluating images another environment.
At the end of the day radiology is more an art than a science, so the training data may well be faulty. Krupinski (2010) wrote in an interesting paper [1]:
"Medical images need to be interpreted because they are not self-explanatory... In radiology alone, estimates suggest that, in some areas, there may be up to a 30% miss rate and an equally high false positive rate ... interpretation errors can be caused by a host of psychophysical processes ... radiologists are less accurate after a day of reading diagnostic images and that their ability to focus on the display screen is reduced because of myopia. "
I would hope datasets included a substantial amount of images that were originally mis-classified as a human.
Medical data for AI training is almost always sources in some more or less shady country because they lack any privacy regulations. It's then annotated by a hoard of cheap workers who may or may not have advanced medical training.
Even "normal medicine" is extremely biased towards male people fitting inside the norm which is why a lot of things are not detected early enough in women or in people who do not match that norm.
Next thing: Doctors often think that their annotations are the absolute gold standard but they don't necessarily know everything that is in an X-Ray or an MRI.
A few years ago we tried to build synthetic data for this exact purpose by simulating medical images for 3D body models with different diseases and nobody we talked to cared about it, because "we have good data".
Sadly I'd say that people are no different.
> it's trained on horrible cesspools of ....
So it's really not the future of AI we should be worrying about...
My conspiracy: With massive medical data, ML/AI would have been 'discovered'/built sooner. Limiting the data makes it so only a few people can be specialists under the supervision of medical cartels.
This isn't a fully settled debate (including your salary example) so you can't just assume one side is right and argue it as if it's some unquestionable human right.
However, most of that data is useless for research purposes. Even if the format complies with industry standards the quality is often bad with many data elements lacking consistent coding. You can't just feed clinical data from a bunch of different random sources into a research project and expect to get accurate results: it's a garbage in / garbage out issue. That's why most clinical research studies involve just a few provider organizations so that the researchers can properly configure the systems and train the clinicians on consistent data entry.
https://privacyruleandresearch.nih.gov/pr_08.asp
https://www.hhs.gov/hipaa/for-professionals/privacy/index.ht...
They call demographics like age and sex "shortcuts" but I find this to be a frustrating term since it seems to obscure what's happening under the hood. (They cite many papers using the same word, so I'm not blaming them for this usage.) Men are typically larger; old bones do not look like young bones. There is plenty of biology involved in what they refer to as demographic shortcuts.
I think you could take the same results and say "Models are able to distinguish men from women. For our purposes, it's important that they cannot do this. Therefore, we did XYZ on these weakly labeled public databases." But perhaps that sounds less exciting.
A model overfits if it is unnecessarily COMPLEX for the training data.
If there is bias in the training (and validation and test) data that allows a SIMPLE model to fit the data because of a spurious correlation, that is not overfitting.
More specifically, politically undesirable correlation - as in, "it's there, but its existence upsets some people". It's pretty obvious and self-evident that there are meaningful biological differences related to age, sex, and other demographics. Whether or not they're clinically relevant for a specific diagnosis under question is one thing, but they are clinically relevant for great many diagnoses; trying to "de-bias" reality here will only lead to unnecessary suffering and loss of life.
Now you can say that this is perfectly fine and represents the most likely real-world use case. Or you might prefer a model that looks at the image only, with the implicit assumption that this "forbidden knowledge" will be added by human doctors later on in the pipeline. This is beneficial because the "forbidden knowledge", such as whether patients from Hospital A always have bone cancer, might change overnight! Imagine the hospital gets assigned a new name in the system and the prediction is shit now.
This second, "unbiased" AI system will always have a worse performance, because you lobotomize it when you kill the forbidden knowledge with a sledgehammer. This study just showed that "group fairness" is at odds with optimal predictions" and how much it is at odds.
PS: You might even prefer a society where everyone is worse off, but every protected group is equally bad off. You'd also ban the humans from applying the forbidden knowledge. Whether that is desirable, is, of course, out of the scope of the paper.
These models are only looking at the images. They are inferring demographics.
Your knowledge guides, but it also doesn't (or shouldn't) blind you.
That the model can determine biological sex from X-rays wouldn't be an issue if it never shortcuts the diagnostic process by using biological sex in place of meaningful diagnostic data. I would not like a model to ignore a melanoma in my chest scan because it can deduce that I was born male and my risk of breast cancer is quite low.
The idea of penalising a model which takes such biological shortcuts (because its subgroup accuracy gets worse) seems like a good solution, and it's cool that the approach works in TFA.
There are plenty of tools that have been developed to enforce that models perform in a manner that is unbiased across some dimension (e.g., sex, or hospital, etc). (For example unsupervised domain adaptation.) I think that splitting the field with jargon makes it more difficult to follow the breadth of the field.
"I think the main takeaways are, first, you should thoroughly evaluate any external models on your own data because any fairness guarantees that model developers provide on their training data may not transfer to your population. Second, whenever sufficient data is available, you should train models on your own data," says Haoran Zhang, an MIT graduate student and one of the lead authors of the new paper. ```
This is just overfitting. Why are they training whole models on only one hospital worth of data when they appear to have access to five? They should be training on all of the data in the world they can get their hands on then maybe fine tuning on their specific hospital (maybe they have higher quality outcomes data that verifies the readings) if there are still accuracy issues. The last five years have taught us that gobbling up everything (even if it's not the best quality) is the way.
Like the time of death after the data was collected.
If a model could with a high accuracy predict, that a patient will die within X days (without proper treatment), it will be already very valuable.
Second, as Sora has shown, going multi model can have amazing benefits.
Get a breath analysis of the patient, get a video, get a sound recording, get an MRI, get a CT, get a full blood sample and then let the model do its pattern finding magic.
It's not a lack of "fairness", it's just a lack of accuracy.
Imagine that you train a model to find roofing issues or subsidence or something from aerial imagery. Maybe it performs better on Victorian terraces, because there are lots of those in the UK.
Would you call it unfair because it doesn't do so well on thatched roof properties? No, it's just inaccurate, calling it unfair is a value judgement.
Bias is better because it at least has a statistical basis but fairness is, well.. inaccurate...
Although now that I think of it, there could be a "fairness" angle regarding how many training data are obtained from each population group.
But if that's what the author meant by fairness, it wasn't clear.
On top of that fairness in the general public is undefined as well. I once wrote on the many facets just concerning fair salaries here: https://www.amazingcto.com/fair-remote-developer-salaries/
You would also need to assume there no overlap in the distribution in the training data for each strata of training data. I would imagine that if you taught someone to detect cancer in x-rays, and only used images of people aged 80+, you might still have some success (>0.5) at detecting cancer in people aged 20-30.
The greater issue is that, for privacy and other reasons, there are many demographics that are not represented at all, or under represented, in medical image datasets.
(Although apparently that's sometimes debated [0].)
I guess that these factors are harder to prove though.
If you ask me, it's probably more effective to compensate for the model's learned racial bias using weights derived from the model outputs via statistical analysis.
Suppose that older people are more likely to get cancer, so older people with cancer are more represented in the training data. Then we discover that it has a 20% false negative rate for older people but a 35% false negative rate for younger people, i.e. it has a better handle on what cancer looks like in older people. This is fairly intrinsic, the only real way to fix it is to provide it with more samples of images of younger people with cancer, which may not be available.
But you can also "fix" it by raising the false negative rate for older people to 35%. Doing this is idiotic and should not be done, but it allows people to claim that the model is now "fair", and so they will have the incentive to do things like that as long as we continue to demand "fairness" above other considerations.
Moreover, you're likely to see similar effects with all kinds of other groups. Maybe farm workers are more likely to get skin cancer because they spend more time in the sun, and you could see the same kind of disparity in its ability to detect cancer in farm workers because the effects of doing farm work also have other long-term physical correlates that show up in the image. Probably nobody is studying this because farm workers are not a protected class, but it's just a thing that happens whenever you divide the population into distinct groups. That doesn't mean there is inherently anything to be done about it. It's a facet of what data is available.
FYI a "protected class" is an attribute category, not an attribute value. In your example it would be something like "career" or "employment type", rather than "farmer".
[0] https://content.next.westlaw.com/practical-law/document/Ibb0...