Their results seem solid, and clear, to me.
When using a chest x-ray to look for pulmonary edema, for instance, I would be unsurprised if breast tissue (of any quantity) and in particular denser breast tissue would make the diagnosis of pulmonary edema more difficult from the image alone.
Also, you seem to have conflated a few things in your second sentence. Deep in the article, they did have radiologists try to guess demographic attributes by looking at the x-ray images. They were pretty good at guessing female/male (unsurprising) and were not really able to guess age or race. So I'm super interested in how the AI model was able to be better at that than the human radiologists.
For example, a couple years ago there was a statistical model made which could fairly accurately predict (iirc >80%) the gender of a person based on a picture of their iris. At the time we didn’t know there was a visible iris difference between genders, but a statistical model found one.
That’s kind of the whole point of statistical classification models. Feed in a ton of data and the model will discover the differentiating features.
Put another way, If we knew all the possible differences between someone with cancer and without, we wouldn’t need statistical models at all, we could just automate the diagnosis.
We don’t know the indicators that we don’t know, so we don’t know if some possible indicators show up or don’t show up in a given group of people.
That is the danger of wholly relying on statistical models.
Black women experience worse outcomes and are diagnosed with more severe forms of breast cancer than white women.
Cancer is not just one disease. Its progression will vary depending on type. If the AI is trained on only some strains of cancer, eg those traditionally found in white women in early detection scenarios, it might not generalize to other cancer types.
So yes, to your genuine question, medical imaging of cancer can vary depending on ethnicity because different cancers can vary between genetic backgrounds. Ideally there would be sufficient training data across the populations, but there isn't because of historical race bias. (Among other reasons.)