I suspect we massively underestimate the amount of misdiagnosis due to incorrect analysis of data using fairly naive medical mental models of disease.
I suspect we massively underestimate the amount of misdiagnosis due to incorrect analysis of data using fairly naive medical mental models of disease.
Personal story: I was diagnosed with a rare genetic disease in 2019. If I ran the symptoms through a ML gauntlet, I would be sure they would cancel each other out or make little sense. Chest CT (clean), fever (high), TB test (negative), latent TB marker (positive), vision difficulty (Nothing unusual yet), edema in eye socket (yes), WBC count (normal), tumors (none), hormones (normal) & retina images (severely abnormal)
My condition was zeroed in within 5 minutes of a visitation to a top retina specialist, after regular opthalmologists were in a fix about two conflicting conditions. This was differential diagnosis based even though genetic assay hadn't returned yet, which also later came in favor. I cannot overemphasize enough how good human brain is in recalling information & connecting the sparse dots to logical conclusions
(I am one of 0.003% unlucky ones among all opthalmological cases & the only active patient with that affliction in one of the busiest hospitals in the country. My data is part of the 36 people in a NIH study & opthalmo residents are routinely called in to see me as case study when I go for follow up quarterly).
How many other people with the condition were misidentified?
I only say this because of a family member with a rare genetic condition. For years they were told it was something else, or told 'it was in their head'. The family member started a journal of their medical conditions and experiences that was detailed then brought that to their PC which whom then sent them to a specialist, this specialist wasn't sure and sent them to another specialist that had a 3 month wait. After 5+ years of living with increasing severity of the condition it was identified.
So, just saying, it's as much likely that the condition was identified because you kept a detailed list (on paper or in your mind) of the aliments and presented them in a manner that helped with the final diagnosis.
2 opthalmo, 1 internal medicine, 1 retina super-specialist & finally someone from USC Davey
> How many other people with the condition were misidentified?
Historical data: I don't know. It is fairly divided between two types, one being zoonotic & other to IL2 gene. I am told this distinction of pathways was identified in 2007.
> [..] you kept a detailed list (on paper or in your mind) of the aliments and presented them in a manner that helped with the final diagnosis.
I might have been a better informed patient but I went with a complaint of pink eye, flu & mild light sensitivity. Never imagined that visit would change my life forever. Thank you though, for expressing your concern & support
Human minds can be really good at diagnostics and still fail sometimes when faced with very difficult cases.
In my experience, ML would just classify everything as a very common disease and people would call it a success because it has an 80% effectiveness rate.
The problem that needs to be solved is a case like your example, not diagnosing the common cold.
One of the points taken should be that either via ML or human diagnostic is that these rare problems are either not diagnosed for long periods of time reducing quality of life or diagnosed posthumously.
The reduction of these measures are what we should use when making meat vs machine efficiency correlations.
However, there's a lot that isn't covered with data. The "middle of the scale", the "almost but not quite there", the "this is weird"... Doctors are good at that, through experience, and those are the difficult cases. Those are the ones where ML will not only likely fail, but won't even explain why it fails. We're talking about human lives here. If anything, I think software engineers massively overestimate the performance of ML and underestimate doctors.
I think it's probably going to be a long time before models only using quantifiable measurements can even meet the performance of top doctors. I can't recommend enough that someone experiencing issues doctor-shop if they haven't gotten a well-explained diagnosis from their current doctor.
But I'm very curious how good one has to be in order to be better than a below-average doctor, or a 50th-percentile doctor, or a 75th...
But I also think there may be weird failure modes similar to today's not-fully-self-driving cars along the lines of "if even the 75th-percentile-doctor uses the tool and sees an output that stops them from asking a question they otherwise might have, can it hurt things too?"
In dermatology, on which I was working, models were better (at detecting skin cancers) than 52% of the GPs, going by just images. In a famous Nature paper by Esteva et al., the TPR was at 74% for detecting Melanomas. There is a catch which probably got underreported (The skin cancer positivity rate was strongly correlated to clinical markings in photos. Their models didn't do quite as well when 'clean', holdout data were used).
But the nature of information in all these models were skin deep (pun intended). They were designed with a calibrated objective in place unlike how we approach clinical diagnostics as open ended problems for the doctors.
When they went back and tested with clean images that didn't basically have the "im a positive cuz I have this label over here in the margins", the hit rate dropped below that of humans.
It was an article with anecdotes about some of the hospice cats that seemingly are able to detect when a patient is about to die. Entirely possible as they have a sense of smell and patient tumors likely giving out detectable odors.
Nonetheless, the ML model & the cat were similarly inscrutable.
From Andre Esteva's paper:
And yes I'm a physician and MLE. So i understand both worlds clearly
It is kind of funny to see comments complain about the lack of perfect sensitivity and specificity of their physicians.
We complain about the same thing from the the various ML techniques in radiology which currently are pitiful and a gigantic waste of time and money. When I went into rads I was pretty worried about ML - not anymore.
I’m hoping this upcoming recession will dry out a lot of institutional use of ML. In radiology it’s not that helpful and there’s no technical fee increase for it. But you can advertise with it i guess? Commercials with lasers, robots, and AI with pleasant voice overs about cutting edge techniques and getting the care you need in the 21st century and blah blah blah
I suspect software engineers massively underestimate the value of skills outside their domain.
The same applies not just for software devs, but for every other domain as well.