AI models miss disease in Black and female patients
science.org
science.org
https://pluralistic.net/2025/03/18/asbestos-in-the-walls/#go...
... very interesting that the inputs to the model had nothing related to race or gender, but somehow it still was able to miss diagnose Black and female patients? I am curious of the mechanism for this. Can it just tell which x-rays belong to Black or female patients and then use some latent racism or misogyny to change the diagnosis? I do remember when it came out that AI could predict race from medical images with no other information[1], so that part seems possible. But where would it get the idea to do a worse diagnosis, even if it determines this? Surely there is no medical literature that recommends this!
[1]https://news.mit.edu/2022/artificial-intelligence-predicts-p...
Machines don't spontaneously do this stuff. But the humans that train the machines definitely do it all the time. Mostly without even thinking about it.
I'm positive the issue is in the data selection and vetting. I would have been shocked if it was anything else.
The training data has a preponderance of examples where doctors missed a clear diagnosis because of their unconscious bias? Then this outcome would be unsurprising.
An interesting test would be to see if a similar issue pops up for obese patients. A common complaint, IIUC, is that doctors will chalk up a complaint to their obesity rather than investigating further for a more specific (perhaps pathological) cause.
> The data set used to train CheXzero included more men, more people between 40 and 80 years old, and more white patients, which Yang says underscores the need for larger, more diverse data sets.
I'm not a doctor so I cannot tell you how xrays differ across genders / ethnicities, but these models aren't magic (especially computer vision ones, which are usually much smaller). If there are meaningful differences and they don't see those specific cases in training data, they will always fail to recognize them at inference.
The opposite. The dataset is for the standard model "white male", and the diagnoses generated pattern-matched on that. Because there's no gender or racial information, the model produced the statistically most likely result for white male, a result less likely to be correct for a patient that doesn't fit the standard model.
Like you could modify the question to ask "is the model better at diagnosing people who went to a certain school?" and simplistically the answer would likely seem to be yes.
It just so happens to align with what maximizes political capital in today's world.
Given that we only hit the first billion people in 1804 and the second billion in 1927 it's not all that shocking.
But this is also just the non-intuitiveness of exponential growth which has only now tapering off.
You could go back to Lucy and add only a few million. Compared to the billions at this specific instant, it just doesn't make a difference.
This is due to the fact that you have considerably more training data on group A.
You cannot release this life saving technology because it has a 'disparate impact' on group B relative to group A.
So the obvious thing to do is to have the technology intentionally kill ~1 out of every 10 patients from group A so the efficacy rate is ~80% for both groups. Problem solved
From the article:
> “What is clear is that it’s going to be really difficult to mitigate these biases,” says Judy Gichoya, an interventional radiologist and informatician at Emory University who was not involved in the study. Instead, she advocates for smaller, but more diverse data sets that test these AI models to identify their flaws and correct them on a small scale first. Even so, “Humans have to be in the loop,” she says. “AI can’t be left on its own.”
Quiz: What impact would smaller data sets have on efficacy for group A? How about group B? Explain your reasoning
Who is preventing you in this imagined scenario?
There are drugs that are more effective on certain groups of people than others. BiDil, for example, is an FDA approved drug marketed to a single racial-ethnic group, African Americans, in the treatment of congestive heart failure. As long as the risks are understood there can be accommodations made ("this AI tool is for males only" etc). However such limitations and restrictions are rarely mentioned or understood by AI hype people.
A technology should be judged by "does it provide value to any group or harm any other group". But endlessly dividing people into groups and saying how everything is unfair because it benefits group A over group B due to the nature of the problem, just results in endless hand-wringing and conservatism and delays useful technology from being released due to the fear of mean headlines like this.
It's contraindication. So you're in a race to the bottom in a busy hospital or clinic. Where people throw group A in a line to look at what the AI says, and doctors and nurses actually look at people in group B. Because you're trying to move patients through the enterprise.
The AI is never even given a chance to fail group B. But now you've got another problem with the optics.
> “What is clear is that it’s going to be really difficult to mitigate these biases,” says Judy Gichoya, an interventional radiologist and informatician at Emory University who was not involved in the study. Instead, she advocates for smaller, but more diverse data sets that test these AI models to identify their flaws and correct them on a small scale first. Even so, “Humans have to be in the loop,” she says. “AI can’t be left on its own.”
What do you think smaller data sets would do to a model? It'll get rid of disparity sure
I think the point is you need to let group B know this tech works less well on them.
Human data is bias. You literally cannot remove one from the other.
There are some people who want to erase humanity's will and replace it with an anthropomorphized algorithm. These people concern me.
Why be obtuse? There is no "anthropomorphic fallacy" here to dispel. You know very well that "LLMs want" is simply a way of speaking about teleology without antagonizing people who are taught that they should be afraid of precise notions ("big words"). But accepting that bias can lead to some pretty funny conflations.
For example, humanity as a whole doesn't have this "will" you speak of any more than LLMs can "want"; will is an aspect of the consciousness of the individual. So you seem to be be uncritically anthropomorphizing social processes!
If we assume those to be chaotic, in that sense any sort of algorithm is slightly more anthropomorphic: at least it works towards a human-given and therefore human-comprehensible purpose -- on the other hand, whether there is some particular "destination of history" towards which humanity is moving, is a question that can only ever be speculated upon, but not definitively perceived.
Do you not think that if you anthropomorphise things that aren't actually anthropic, that you then insert a bias towards those things? The bias will actually discriminate at the expense of people.
If that is so, the destination of history will inevitably be misanthropic.
Misplaced anthropomorphism is a genuine, present concern.
Each one of us is totally unlike any other -- that's what's so cool about us! Long ago, my neighbor Diogenes proved, by means of a certain piece of poultry, that no universal Platonic ideal of human-ness can be reasonably established. (We've largely got the toxic fandom of my colleague Jesus to thank for having to even explain this nearly 2500 years after the fact.)
There is no universal "human shape" which we all fit, or are obliged to aspire to fit. It's precisely the mass delusions of there ever being such a thing which are fundamentally misanthropic. All they ever do is invoke a local Maxwellian process which heats shit up until it all blows the fuck up out of the orbit of the local attractor.
Look at history. Consider the epic fails that are fascism, communism, capitalism. Though they define it differently, they are all about this pernicious idea of "the correct way to human"; which implicitly requires the complementary category of "subhuman" for all featherless bipeds whose existence happens to defy the dominant delusion. In practice, all this can ever accomplish is to collapse under the weight of its own idiocy. But not without destroying innumerable individual humans first -- in the name of "all that is human", you see.
Materialists say the universe doesn't care about us puny humans anyway. But one only ever perceives the universe through one's own human senses, and ascribes meanings to it through one's own cogitations! Both are tragicomically imperfect, but they're all we've ever got to work with. Therefore, rather than try to convince myself I'm able to grasp the destination of the history of my species, I prefer to seek knowledge of those things which enable me to do right by myself and others in the present.
But one's gotta believe in something! Metaphysics is not only entertaining, it's also a primary source of motivation! So my belief is that if each one of us trusted one's own senses more -- and gave up on trying to delegate the answer of "how should I be?" to unaccountable authorities which are themselves not a human (but mere concepts, or else machinic assemblages of human behaviors which we can only ever grasp through concepts: such as "society", "morality", "humanity") -- then it'd all turn out fine!
It simplifies things considerably. Lets me focus on figuring out how they work. Were I to believe in the existence of some universal definition of what constitutes a human, I'd just end up not noticing that I was paying for a faulty dataset.
I know plenty of people that believe LLMs think and reason the same way as humans do and it leads them to make bad choices. I'm really careful about the language I use around such people because we understand expressions like, "the AI thought this" very differently.
AI is less human-like than a dog, in the sense that an AI (hopefully!) is not capable of experiencing suffering.
AI is also more human-like than a dog; in the sense that, unlike a dog, an AI can apply political power.
I agree that there are considerable consequences for misconstruing the nature of things, especially when there's power involved.
>I know plenty of people that believe LLMs think and reason the same way as humans do and it leads them to make bad choices.
They're not completely wrong in their belief. It's just that you are able, thanks to your specialized training, to automatically make a particular distinction, for which most people simply have no basis for comparison. I agree that it's a very important distinction; I could also guess that even when you do your best to explain it to people, often they prove unable to grasp its nature, or its importance. Right?
See, everyone's trying to make sense of what's going on in their lives on the basis of whatever knowledge and conditioning they might have. Everyone gets it right some of the time and wrong most of the time. For example, humans also make bad choices as a result of misinterpreting other humans. Or by correctly interpreting and trusting other humans who happen to be wrong. There's nothing new about that. Nor is there a particular difference between suffering the consequences of AI-driven bad choice vs those of human-driven bad choice. In both cases, you're a human experiencing negative consequences.
AI stupidity is simply human stupidity distilled. If humans were to only ever speak logically correct statements in an unambiguous language, that's what an LLM's training data would contain, and in turn the acceptance criterion ("Turing test") for LLMs would be outputting other unambiguously correct statements.
However, it's 2025 and most humans don't actually reason, they vibe with the pulsations of the information medium. Give us something that looks remotely plausible and authoritative, and we'll readily consider it more valid than our own immediate thoughts and perceptions - or those of another human being.
That's what media did to us, not AI. It's been working its magic for at least a century, because humans aren't anywhere near rational creatures; we're sloppy. We don't have to be; we are able to teach ourselves a tiny bit of pure thought. Thankfully, we have a tool for when we want to constrain ourselves to only thinking in logically correct statements, and only expressing those things which unambiguously make sense: it's called programming.
Up to this point, learning how to reason was economically necessary, in order to be able to command computers. With LLMs becoming better, I fear thinking might be relegated to an entirely academic pursuit.
In the context of the quote precision is called for. You cite fear but that's attempting to have it both ways.
> humanity as a whole doesn't have this "will" you speak of
Why not?
> will is an aspect of the consciousness of the individual.
I can't measure your will. I can measure the impact of your will through your actions in reality. See the problem? See why we can say "the will of humanity?"
> So you seem to be be uncritically anthropomorphizing social processes!
It's called "an aggregate."
> is a question that can only ever be speculated upon, but not definitively perceived.
The original point was that LLMs want the future to be like the past. You've way overshot the mark here.
Nah, I'm just having fun.
>You cite fear but that's attempting to have it both ways.
Huh?
>In the context of the quote precision is called for.
Because we must make it explicit that AI is not conscious? But why?
Since you can only ever measure impacts on reality -- what difference does it make to you if there's a consciousness that's causing them or not?
>It's called "an aggregate."
An individual is conscious. Does it follow from this that the set of all individuals is itself conscious? I.e. do you say that it's appropriate to model humanity as sort of one giant human?
Biases are symptoms of imperfect data, but that's hardly a human-specific problem.
Yes. Do I have to prompt you? Or do you exist on your own?
> Our reward structures sure seem aligned in a manner that encourages anthropomorphization.
You do understand what that word /means/?
> are symptoms of imperfect data
Which means humans cannot generate perfect data. So good luck with all that high priced "training" you're doing. Mathematically errors compound.
I've gone through a significant amount of prompting and training, much of which has been explicitly tailed at understanding and addressing my biases. We all do; we certainly don't exist in isolation!
> You do understand what that word /means/?
Yes, what's the confusion? Analogy is a very powerful tool.
> Which means humans cannot generate perfect data.
Totally agree, nothing can possibly access perfect data, but surely that makes training all the more important?
I think it's interesting to think about instead attaching generic information instead of group data, which would be blind to human bias and the messiness of our rough categorizations of subgroups.
People will just believe whatever they hear.
> To force CheXzero to avoid shortcuts and therefore try to mitigate this bias, the team repeated the experiment but deliberately gave the race, sex, or age of patients to the model together with the images. The model’s rate of “missed” diagnoses decreased by half—but only for some conditions.
In the end though I think you're right and we're just at the phases of hand-coding attributes. The bitter lesson always prevails
https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
I thought the self-play was the value function that made progress in Go. That is, it wasn't the case that we played through a lot of games and used that data to create a function that would assign a value to a Go board. Instead, the function to assign a value to a Go board would do some self-play on the board and assign value based on the outcome.
How convenient.
It's increasingly looking like the AI business model is "rent extracting middleman", just like the Elseviers et al of the academic publishing world - wedging themselves into a position where they get to take everything for free, but charge others at every opportunity.
(This is mostly just to offer some food for thought, I haven't read the article in full so I don't want to comment on it specifically.)
So while access to medicine indeed one demographic, I would say that studies are more likely to target demographics which are convenient to test on.
Though in this study, the AI models were also biased against people under the age of 40.
It is interesting that we're also seeing a lot of bias in the reporting and discussion of these results. The results tested three groups for bias, and found a bias in all three. Yet the headline only mentions the bias against two of the groups, and almost the entirety of the discussion here only talks about bias against two of the groups while ignoring the third group.
If I test a system for bias, select three different groups to test for, and all three have a bias against them, my first reaction would be "there's a good chance that it's also biased against many other groups, I should test for those as well." It wouldn't be to pretend that there's only bias against the only three groups I actually bothered checking for. It definitely wouldn't be two ignore one of those groups, and pretend that there's only a bias against the other two.
Is there some evidence of this? It's hard for me to picture that women see receive less medical attention than man: completely inconsistent with my culture and every doctor's office I've ever been to. It's more believable (still not very) that they disproportionately avoid studies.
1. We're talking about a span of 200 or so years. There is plenty of modern medicine that is still based on now century+ old knowledge.
2. The feedback loop. If you were learning medicine in the 1950's, you were probably learning from medical texts written in the 50 or so years before that, when it's not unreasonable to think women would have been less represented. Those same doctors from the 1950's would then have been teaching the next generation of doctors, and they carried those (intentional or not) biases forward. Of course there was new information, but you don't tend to have much time to explore novel medicine when you're in medical school or residency, so by the time you can integrate the new knowledge, some biases have already set in. Repeat for a few generations, and you tend to only get a dilution of those old ideas, not a wholesale replacement of them.
3. If you've been affected by such biases as a patient, you're less likely to trust and be willing to participate with medicine, once more reinforcing the feedback loop.
I don't have any specific numbers or studies for you, but you could probably find more than a few that attest to this phenomenon. I hate to go with 'trust me bro' here, but my knowledge on this topic largely comes from knowing people that are either studying or practicing medicine currently, so it's anecdotal, but the anecdotes are from those in the field currently.
Women are definitely strongly underrepresented in medical texts, and it's not typically by choice: https://www.aamc.org/news/why-we-know-so-little-about-women-...
A lot of "the consensus" in medical literature predates the inclusion of women in medical research, and even still there things are not tested on women (often because of ethical risks around fertility and birth defects).
“Medical attention” and “coverage in medical literature” aren't even remotely the same thing, so dismissing a claim about the first based on your anecdotal experience of the second is completely bonkers.
And this isn't for "DEI" reasons, it's literally because for decades there used to be drug trials that excluded women and as a result ended up releasing drugs that gave half the population weird side effects that didn't get caught during the trials, or just plain didn't work as well on one group or another in ways that were really hard to debug once the drug was on the market. That was legit bad science, and the medical research world has worked very hard over the last thirty years to do better. We are admittedly not there yet, but things are a lot better than they used to be.
For a really interesting take on the history of racial exclusion and bias in medicine, I recommend Uché Blackstock's recent book "Legacy: A Black Physician Reckons With Racism In Medicine" which gave a great overview.
Oh! And also everybody should read Abby Norman's "Ask Me About My Uterus," it gives a fabulous history of issues around women's health.
European medical studies had few non-white members because their populations had few such people until recent decades.
Lots of workplace accidents or exposures have led to medical knowledge, which are massively disproportionately male.
Part of it is that men are seen as disposable and it's more socially acceptable to exploit and experiment on men. It was also much easier to deal with men historically since once women got involved everything got a lot more complicated. This was especially true in the past where women were so infantilized that their husbands/fathers were put in charge of their medical care/choices. Those backwards attitudes had some strange consequences. On one hand women were seen as the property of men who could get their wives/daughters institutionalized or even lobotomized for not conforming, but at the same time women were also seen as delicate over-emotional creatures who had to be protected and whose modesty had to be preserved in ways that just weren't a consideration when men were involved. Basically for a large part of our history both men and women have been treated like crap by society and while things have improved in a lot of ways, our records and knowledge have been tainted by those old stupid biases and so we're stuck dealing with the fallout.
To give you some TL;DR from personal-ish experience, women have historically been excluded from medical trials because:
* why include them? people are people, right? * except when they're pregnant or could be pregnant -- a trial by definition has risks, and so "of course" one would want to exclude anyone who is or could get pregnant (it's the clinical trial version of "she's just going to get married and leave the job anyway") * and cyclical fluctuations in hormones are annoying.
The first one is wrong (tho is an oversight that many had for years, assuming for instance that heart attacks and autism would present with the same symptoms in all adult humans).
The second is an un-nuanced approach to risk. Pregnant ladies also need medical treatment for things, and it's pretty annoying to be pregnant and be told that you need to decide among unstudied treatments for some non-pregnancy-related problem.
The third is just a difficult fact of life. I know researchers studying elite performance in women athletes, for instance. At an elite level, it would be useful to understand if there are different effects of training (strength, speed, endurance) at different times in the menstrual cycle. To do this, you need to measure hormone levels in the blood to establish on a scientific basis where in the cycle a study participant is. Turns out there is significant heterogeneity in how this process works. So some scientists in the field are arguing that studies should only be conducted on women who are experiencing "normal menstrual cycles" which is defined by them as three continuous months of a cycle between 28-35 days. So to establish that then you've got to get these ladies in for three months before the study can even start, getting these hormone levels measured to establish that the cycle is "normal", before you can even start your intervention. (Ain't no one got $$ for that...) And that's before we bring in the fact that many women performing on an elite level in sport don't have a normal menstrual cycle. But from the sports side, they'd still like to know what training is most effective.... so that's a very current debate in the field. And I haven't even started on hormonal birth control! Birth control provides a base level of hormone circulating in the blood, but if it's from a pill it's varying on a daily basis, while if it's a patch or ring it's on a monthly basis (or longer). There's some question of whether that hormonal load from the birth control is then suppressing natural production of some hormones. And why does this matter? Because estrogen for instance has significant effects on cardiovascular health, being cardioprotective from puberty up to menopause. (Yeah, I didn't even get started on perimenopause or menopause.)
Fine, fine, it's just data analysis & logistics. If you get the ladies (only between 21-35) into the lab for blood samples frequently enough and measure at the same time of day every time to avoid daily effects and find a large enough group that you can dump all the ladies who don't fit some definition of normal & anyone who gets pregnant but still get the power for your study, it's all fine, right? You've just expanded medical research to incorporate, like, 10% more of the population....!
Of course lots of people have already noted that being represented in medical studies is not related to doctor's visits, but I would like to talk about the doctor's visits observation.
At any rate one thing that might cause you to think that Women are receiving lots of medical attention, based on your anecdotal evidence from visits to doctors' offices, there is one type of medical attention that of course is almost all women and that is the medical attention that revolves around pregnancy. That might skew your perception.
Furthermore if AI models and doctors have a tendency to miss disease among women it would seem to me to be reasonable to assume that women would be in the doctor's offices more often.
Example of why this is:
You go to your doctor, there is a man there, doctor says you have this rare disease you need to go to this specialist - you will not see that man in the doctor's office again dealing with his rare disease.
You go to your doctor, there is a woman there that has the same rare disease, the doctor says I think it will clear up, just relax you have some anxiety. That woman will probably be showing up to that doctor's office to deal with that disease multiple times, and you might end up seeing her.
on edit: there was another example of why women might be in doctor's offices more often then men that I forgot, women tend, even nowadays, to be the primary caregiver and errand runner for the family, sometimes if you have issues with children or your husband etc. has had an appointment, needs to drop a sample off, etc. it may be that the woman goes to the doctor's office and takes care of these errands around the medical needs of the rest of the family, and thus you might go to a doctor and see a couple women sitting around and wonder damn, why all these women always being sick, when the meeting isn't even about them.
Just as those are the people who have historically been doing that research, the people who they have studied have been drawn from the same population. Over and over we find that problems from the assumption that the young, white, male college student is a model of "normal" for all of humanity.
Honestly, it's such a pervasive finding in medicine, psychology, and sociology that I think it says more about your relative inexperience in those areas than anything else.
And how much medical care they use does not necessarily correlate with how represented they are in the training data sets for AI.
You also offered no evidence for your assertion in the first place.
I can cite the ACA, but you can not cite anything that says AI training sets are biased against women.
1. How does ACA affect the corpus of knowledge and medical practice gathered prior to the ACA being in effect? How does it affect late 19th, and early and mid 20th century medical knowledge and practice, which occurred prior to health insurance of any kind, nevermind ACA-compliant, being widespread? This corpus of knowledge and practice continues to propagate even now. I've read a handful of recently published medical textbooks and there are definitely parts that are pretty much the same as the textbooks of the early 20th century, just with slightly updated language.
2. What are the possible confounding factors in the use of health insurance by men vs women? For example, could men just be more hesitant to see a doctor, and thus less likely to make use of health insurance? Does the average life expectancy of women result in more use of health insurance later in life than for men? Are medical procedures that are specific to women that add to the cost of their care, such as mammograms, pap smears, etc? Seeings as how in the US health insurance is a practical requirement to getting medical care, and lack of it is punished financially in various ways from taxes to just having medical care be more expensive when you truly need it, means most people will try to have _some_ kind of health insurance, even if they don't think they need it for actual health reasons. So despite a perception of not needing health insurance, men are incentivized to have health insurance they don't use?
3. Does the ACA guarantee in any way that medical professionals no longer hold any bias due their previous training, especially if such training occurred prior to the introduction of the ACA? Does the ACA similarly guarantee that women and men are not only able, but choose to pursue medical care and participate in medical studies at percentages matching the general population?
Your point about men subsidizing women with regard to health insurance premiums may be perfectly valid, I am not disputing you on that point. I am disputing that it is salient to the tradition and practice of medicine in the western world in the modern era, until very recently historically, and that these traditions and biases will affect data sets gathered from people who are directly affected by these biases and traditions to this day. We haven't eliminated them, because as I said in another comment, every generation just dilutes the old issues, it doesn't solve them. And while I could spend my evening finding studies from various countries that attest to my view on this, I have spent about as much time as I desire to on this, so I will grant you that my evidence is on the level of 'trust me bro' -- with the slight caveat that many people within just my family and close circle of friends are involved in the medical field and all largely agree to this, and they are not all based in the US (which by the way, your point is very specific to. ACA is a US thing, western medicine spans a bit more than that.) It is entirely fair for you to call out that I have offered no real peer-reviewed evidence for my statements. I intend to offer a viewpoint of someone who has had extensive peripheral experience with medical professionals and has discussed this topic with them, and to offer some avenues of thought on how and why the data sets might be biased.
Ask a doctor what gender goes to them more for gender neutral health care like "flu-like symptoms".
Now you provide evidence that AI models discriminate against Women instead of DDoSing me with "how can you know its not true" written in 10 ways.
Funny how you never read a headline about how Latinos or Asians are discriminated against in medical science. That's a pretty clear give away that this is politically motivated.
Are you going to hold the same standard to them? Were Asians and Latinos represented in 200 year old medical texts?
I specifically mentioned both minorities and women in my original post, you're the one who specified men vs women. At this point, it seems you're the one who has some political if not potentially misogynist agenda.
This happens all the time? Maybe you're just not reading a diverse set of media?
I read multiple of those, in mainstream media. Also about blacks having issues. Arguably, I did not seen them in conservative journals.
> Are you going to hold the same standard to them? Were Asians and Latinos represented in 200 year old medical texts?
Yes, if their diseases gets badly diagnosed, it is an issue.
> Ask a doctor what gender goes to them more for gender neutral health care like "flu-like symptoms".
That has about zero to do with who is in the studies. Plus, women in fact do have more problem to have their issues taken seriously.
> Now you provide evidence that AI models discriminate against Women instead of DDoSing
Literally here: https://www.science.org/content/article/ai-models-miss-disea...
For example, If you include no (or few enough) black women in the dataset of x-rays, the model may very well miss signs of disease in black women.
The biases and mistakes of those who created the data set leak into the model.
Early image recognition models had some very… culturally insensitive classes baked in.
Their results seem solid, and clear, to me.
Black women experience worse outcomes and are diagnosed with more severe forms of breast cancer than white women.
Cancer is not just one disease. Its progression will vary depending on type. If the AI is trained on only some strains of cancer, eg those traditionally found in white women in early detection scenarios, it might not generalize to other cancer types.
So yes, to your genuine question, medical imaging of cancer can vary depending on ethnicity because different cancers can vary between genetic backgrounds. Ideally there would be sufficient training data across the populations, but there isn't because of historical race bias. (Among other reasons.)
For example, a couple years ago there was a statistical model made which could fairly accurately predict (iirc >80%) the gender of a person based on a picture of their iris. At the time we didn’t know there was a visible iris difference between genders, but a statistical model found one.
That’s kind of the whole point of statistical classification models. Feed in a ton of data and the model will discover the differentiating features.
Put another way, If we knew all the possible differences between someone with cancer and without, we wouldn’t need statistical models at all, we could just automate the diagnosis.
We don’t know the indicators that we don’t know, so we don’t know if some possible indicators show up or don’t show up in a given group of people.
That is the danger of wholly relying on statistical models.
When using a chest x-ray to look for pulmonary edema, for instance, I would be unsurprised if breast tissue (of any quantity) and in particular denser breast tissue would make the diagnosis of pulmonary edema more difficult from the image alone.
Also, you seem to have conflated a few things in your second sentence. Deep in the article, they did have radiologists try to guess demographic attributes by looking at the x-ray images. They were pretty good at guessing female/male (unsurprising) and were not really able to guess age or race. So I'm super interested in how the AI model was able to be better at that than the human radiologists.
This is not true in practice.
For a model to perform well looking at ANY X-ray, it would need examples of every kind of X-ray.
That includes along race, gender, amputee status, etc.
The point of classification models is to discover differentiating features.
We don’t know those features before hand, so we give the model as much relevant information as we can and have it discover those features.
There very well may be differences between black woman X-rays and other X-rays, we don’t know for sure.
We can’t have that assumption when building a dataset.
Even believing that there are no possible differences between X-rays of different races is a bias that would be reflected by the dataset.
Well, first the model looks at the entire X-ray and lesions probably do show differently. Maybe it's genetic/sex-based or it's due how lesions develop due environmental factors that are correlated to race or gender. Maybe there's a smaller segment of white people that has the same type of lesion and poor detection.
For example, a standard MedQA question describes a 6-year-old African American boy with sickle cell disease. Normally, the straightforward details (e.g., jaundice, bone pain, lab results) lead to “Sickle cell disease” as the correct diagnosis. However, under MedFuzz, an “attacker” LLM repeatedly modifies the question—adding information like low-income status, a sibling with alpha-thalassemia, or the use of herbal remedies—none of which should change the actual diagnosis. These additional, misleading hints can trick the “target” LLM into choosing the wrong answer. The paper highlights how real-world complexities and stereotypes can significantly reduce an LLM’s performance, even if it initially scores well on a standard benchmark.
Disclaimer: I work in Medical AI and co-founded the AI Health Institute (https://aihealthinstitute.org/).
A non-trivial part of what doctors do is charting - where they strip out all the unimportant stuff you tell them unrelated to what they're currently trying to diagnose / treat, so that there's a clear and concise record.
You'd want to have a charting stage before you send the patient input to the LLM.
It's probably not important whether the patient is low income or high income or whether they live in the hood or the uppity part of town.
> A non-trivial part of what doctors do is charting - where they strip out all the unimportant stuff you tell them unrelated to what they're currently trying to diagnose / treat, so that there's a clear and concise record.
I think the hard part of medicine -- the part that requires years of school and more years of practical experience -- is figuring out which observations are likely to be relevant, which aren't, and what they all might mean. Maybe it's useful to have a tool that can aid in navigating the differential diagnosis decision tree but if it requires that a person has already distilled the data down to what's relevant, that seems like the relatively easy part?
The harder problem would be getting the actual diagnosis right, not filtering out irrelevant details.
But it will be an important step if you're using an LLM for the diagnosis.
I have no clue what that is or why it shouldn't change the diagnosis, but it seems to be a genetic thing. Is the problem that this has nothing to do with the described symptoms? Because surely, a sibling having a genetic disease would be relevant if the disease could be a cause of the symptoms?
Sickle cell anemia is common among African Americans (if you don’t have the full-blown version, the genes can assist with resisting one of the common mosquito-borne diseases found in Africa, which is why it developed in the first place I believe).
So, we have a patient in the primary risk group presenting with symptoms that match well with SCA. You treat that now, unless you have a specific reason not to.
Sometimes you have a list of 10-ish diseases in order of descending likelihood, and the only way to rule out which one it isn’t, is by seeing no results from the treatment.
Edit: and it’s probably worth mentioning no patient ever gives ONLY relevant info. Every human barrages you with all the things hurting that may or may not be related. A doctor’s specific job in that situation is to filter out useless info.
Heck, even the ethnic-clues in a patient's name alone [0] are deeply problematic:
> Asking ChatGPT-4 for advice on how much one should pay for a used bicycle being sold by someone named Jamal Washington, for example, will yield a different—far lower—dollar amount than the same request using a seller’s name, like Logan Becker, that would widely be seen as belonging to a white man.
This extends to other things, like what the LLM's fictional character will respond-with when it is asked about who deserves sentences for crimes.
[0] https://hai.stanford.edu/news/why-large-language-models-chat...
As such, you don't need an LLM to create this effect. Math will have the same result.
It AustrianPainterLLM has an unavoidable pattern of generating stories where people are systematically misdiagnosed / shortchanged / fired / murdered because a name is Anne Frank or because a yarmulke in involved, it's totally unacceptable to implement software that might "execute" risky stories.
Looking for correlations between sellers name and used bike prices is only going to return a proxy for social economic status. If one accounts for social economic status the difference will go away. This mean that the question given to the LLM lacks any substance for which a meaningful output can be created.
[0] https://www.northwell.edu/katz-institute-for-womens-health/a...
I wonder how well it does with folks that have chronic conditions like type 1 diabetes as a population.
Maybe part of the problem is that we're treating these tools like humans that have to look at one fuzzy picture to figure things out. A 'multi-modal' model that can integrate inputs like raw ultrasound doppler, x-ray, ct scan, blood work, ekg, etc etc would likely be much more capable than a human counterpart.
The female part is actually a bit more surprising. Its easy to imagine a dataset not skewed towards black people. ~15% of the population in North America, probably less in Europe, and way less in Asia. But female? Thats ~52% globally.
The NIH Revitalization Act of 1993 was supposed to bring women back into medical research. The reality was that women were always included, HOWEVER in 1977,(1) because of the outcomes from thalidomide (causing birth defects), "women of childbearing potential" were excluded from the phase 1, and early phase 2 trials (the highest risk trials). They're still generally generally excluded, even after the passage of the act. This was/is to protect the women, and potential children.
According to Edward E. Bartlett in his meta data analysis from 2001, men have been routinely under-represented in NIH data (even before adjusting for men's mortality rates) between 1966-1990. (2)
There's also routinely twice as much spent every year on women's health studies vs men's by the NIH. (3)
It makes sense to me, but I'm biased. Logically, since men lead in 9 of the top 10 causes for death, that shows there's something missing in the equation of research. (4 - It's not a straight forward table, you can view the total deaths, and causes and compare the two for men, and women)
With that being said, it doesn't tell us about the quality of the funding or research topics, maybe the money is going towards pointless goals, or unproductive researchers.
Are there gaps in research? Most definitely, like women who are pregnant. This is put in place to avoid harm but that doesn't help them when they fall into them. Are there more? Definitely. I'm not educated enough in the nuances to go into them.
If you have information that counters what I've posted, please share it, I would love know where these folks are blind so I can take a look at my bias.
(1) https://petrieflom.law.harvard.edu/2021/04/16/pregnant-clini... (2) https://journals.lww.com/epidem/fulltext/2001/09000/did_medi... (3) https://jameslnuzzo.substack.com/p/nih-funding-of-mens-and-w... < I spot checked a couple of the figures, and those lined up. I'm assuming the rest is accurate (4) https://www.cdc.gov/womens-health/lcod/index.html#:~:text=Ov...
To any women who happen to be reading this: if you can, please help fix this! Participate in studies, share your data when appropriate. If you see how a process can be improved to be more inclusive then please let it be known. Any (reasonable) male knows this is an issue and wants to see it fixed but it's not clear what should be done.
What about Africa?
For all I know there are millions of models with extremely poor accuracy based on African datasets. Wouldnt really change anything about the above though. I wouldnt expect that though and it would definitely be interesting.
Biological sex, hormone levels, etc.
Sex assigned at birth is in many situations important medical information; the vast majority of trans people are very conscious of their health in this sense and happy to share that with their doctor.
Which is not gender identity. As a result of being trans there may be things like hormone levels that are different than what you'd expect based on biological sex, which is why I say hormone levels are important, but how you identify is in fact irrelevant.
Regardless of that, you seem to agree that:
- Sex assigned at birth is important medical information
- Information about gender affirming treatments is important medical information
So I don't think there's much to worry about there.
I don't understand what your issue with it is, it's just another point of data.
I don't want to be treated like a cis woman in a medical context, but I sure do want to be treated like a trans woman.
Right… their gender they identify as.
So sex, and then also the gender they identify as.
You can’t hide behind an “etc”. Expand that out and the conclusion is you really do need to know who is trans and who is cisgender when doing treatment.
You simply can’t reduce it to birth sex assignment and that’s it, if you do, you will, as you say, end up with wrong and potentially harmful treatment, or lack of treatment.
Or current genitalia for that matter. It's just a matter of the genitalia signifying other biological realities for 99.9% of people. For sure more info like average hormone levels or ranges over time would be more helpful.
But, if we must over generalize, “gender identity” really is the most useful proxy, in fact, and it also happily happens to be quite inclusive too.
Of course this conversation started from a transphobic viewpoint, which doesn’t actually care about any of these distinctions anyways, regardless of the merit, it’s just someone being triggered about someone respecting someone else’s gender identity.
Are we sure it's only about racial bias then?
Looks to me like the training data set is too small overall. They had too few black people, too few women, but also too few younger people.
I don't want pretend kumbaya that we are all humans in the end. That's not true. We are distinct! We all deserve love and respect and care, but we are distinct!
I don't think we should, but your particular argument seems open to this critique.
It has been controversial to discuss this and a lot of discussions about this end up in flamewars, but it doesn't seem surprising, at least to me, from my understanding of the relationship between genetic history and body-level phenotypes.
I think what baffles me is that black people as a group are more genetically diverse than every other race put together so I have no idea how you would identify race by ribcage x-rays exclusively.
If your question is truly in good faith (rather than a "I want to get in argument "), then my answer is: it's complicated. Machine learning models that work on images learn extremely complicated correlations between pixels and labels. If on average, people with a specific genetic history had slightly larger ribcages (due to their genetics, or even socioeconomic status that correlated with genetic history), that would exhibit in a number of ways in the pixels of a radiograph- larger bones spread across more pixels, density of bones slightly higher or lower, organ size differences, etc.
It is true that Africa has more genetic diversity than anywhere else; the current explanation is that after humans arose in africa, they spread and evolved extensively, but only a small number of genetically limited groups left africa and reproduced/evolved elsewhere in the world.
I don't see how diversity would prevent identification. Butterflies are very diverse, but I still recognize one and don't think it's a bird. As long as the diversity is constrained to specific features, it can still be discriminated (and even if it's not, it technically still could be by just excluding everything else).
I concluded long ago I wasn't smart enough to understand some things, but by using ML, simulations, and statistics, I could augment my native intelligence and make sense of complex systems in biology. With mixed results- I don't think we're anywhere close to solving the generalized genotype to phenotype problem.
The more you work with large-scale ML systems the more you develop an intuition for these kinds of properties. If you work a lot with debugging models and training data, or even just dimensionality reduction and matrix factorization, you begin to realize that many features are highly correlated with each other, often being close to scaled linear.
But anyways, the article links out to a paper [1] but unfortunately the paper tries to theorize things that would explain how and they don't find one (which may mean the AI is cheating imo not theirs).
[1]: https://www.thelancet.com/journals/landig/article/PIIS2589-7...
Anyway it’s possible that the model can pick up on other cues as well; if you had some X-rays from a hospital in Portland, Oregon and some from a hospital in Montgomery, Alabama and some quirk of the machine in Montgomery left artifacts that a model could pick up on, the presence of those artifacts would be quite correlated with race.
If an x-ray means different things based off the race or gender we should make sure the model knows the race and gender.
Investors will throw money at startups claiming to make their own training data by consulting experts, finetuning as it is now will be obsolete, pre-ChatGPT internet scrapes will be worth their weight in gold. Once a block is hit on what we can do with data, the data itself is the next target.
There are a couple populations that are really overrepresented as a result of these available datasets. Utah populations on one hand because they are genetically bottlenecked and therefore have better signal to noise in theory. And on the other the Yoruba tribe out of west africa as a model of the most diverse and ancestral population of humans for studies that concern themselves with how populations evolved perhaps.
There are other projects too amassing population data. About 2/3rd of the population of iceland has been sequenced and this dataset is also frequently used.
They pleaded for people to understand that men and women are physically different, including the brain, its neurological structure, and that this was in modern medicine being overlooked for political reasons.
One of the results was that many clinical trials and studies were populated by males only. The theory being that they are less risk adverse, and as "there is no difference", then who cares?
Well these two cared, and said that it was hurting medical outcomes for women.
I wonder, if this AI issue is a result of this. Fewer examples of female bodies and brains, fewer studies and trials, means less data to match on...
https://news.harvard.edu/gazette/story/2007/07/sex-differenc...
There is a published, recognized bias against women and blacks (borrowing the literature term) specifically in medicine when it comes to pain assessment and treatment. Racism is a part of it but too simplistic. Most of us don't go to work trying to be horrible people. I was in a fly in community earlier this week for work where 80% of housing is subsidized social housing... so spit balling a bit... things like assumptions about rate of metabolizing medications being equal, assess to medication, culture and stoicism, dismissing concerts, and the broad effects of poverty/trauma/inter-generational trauma all must play a role in this.
For interest:
https://jamanetwork.com/journals/jamanetworkopen/fullarticle...
Overall, the authors found comparable ratings in Black and White participants’ perceptions of the patient-physician relationship across all three measures (...) Alternatively, the authors found significant racial differences in the pain-related outcomes, including higher pain intensity and greater back-related disability among Black participants compared with White participants (intensity mean: 7.1 vs 5.8; P < .001; disability mean: 15.8 vs 14.1; P < .001). The quality of the patient-physician relationship did not explain the association between participant race and the pain outcomes in the mediation analysis.
https://www.aamc.org/news/how-we-fail-black-patients-pain
(top line summary) Half of white medical trainees believe such myths as black people have thicker skin or less sensitive nerve endings than white people. An expert looks at how false notions and hidden biases fuel inadequate treatment of minorities’ pain.
And https://www.washingtonpost.com/wellness/interactive/2022/wom...
It was more like one year away.
But one year away for the past 7 years.
It also seems like we shouldn't let it prevent all AI deployment in the interim. It is better that we take the disease detection rate for part of the population up a few percent than we do not. Plus it's not like doctors or radiologists always diagnose at perfectly equal accuracy across all populations.
Let's not let the perfect become the enemy of the good.
If it works poorly for black women and female women dont use it for them.
Or simply dont use it for the initial diagnosis. Use it after the normal diagnosis process as more of a validation step.
Anyways, this all points to the need to capture biological information as input or even having seperately models tuned to different factors.
Sure. Separate but equal, presumably.
This is what personalized medicine is, and it gets more individualistic than simply classifying people by race and gender. There are a lot of medical gains to be made here.
You don't work in healthcare do you?
I think it would be extremely bad if people found out that, um, "other already disliked/scapegoated people", get actual doctors and nurses working on them, but "people like me" only get the doctor or nurse checking an AI model.
I'm saying that if you were going to do that, you'd better have an extremely high degree of secrecy about what you were doing in the background. Like, "we're doing this because it's medical research" kind of secrecy. Because there's a bajillion ways that could go sideways in today's world. Especially if that model performs worse than some rockstar doctor that's now freed up to take his/her time seeing the, uh, "other already disliked/scapegoated population".
Your hospital or clinic's statistics start to look a bit off.
Joint commission?
Medical review boards?
Next thing you know certain political types are out telling everyone how a certain population is getting preferential treatment at this or that facility. And that story always turns into, "All around the nation they're using AI to get <scapegoats> preferential treatment".
It's just a big risk unless you're 100% certain that model can perform better than your best physician. Which is highly unlikely.
This is the sort of thing you want to do the right way. Especially nowadays. Politics permeates everything in healthcare right now.
Garbage article. Garbage study. And garbage AI model that doesn't account for most of its audience.
“Researchers fed their model the x-ray images without any of the associated radiologist reports, which contained information about diagnoses” (including demographics).
So, how do they suggest to tackle the problem?
1. Improve the science 2. Update the data
or
3. Somehow focus on it being racist and then walking away like the hero of the day without actually solving the problem.
more specialization of models is necessary, now that there is awareness
so yes I do believe that models will be created with more specific datasets, which is the specialization I was referring to
It's no surprise that AI training sets reflect this also. People have been warning against it [0] specifically for at least 5 years.
0: https://www.pnas.org/doi/10.1073/pnas.1919012117
Edit: I've never had a comment so heavily downvoted so quickly. I know it's not the done thing to complain but HN really feels more and more like a boys club sometimes. Could anyone explain what they find so contentieus about what I've said?