A.I. Chatbots Defeated Doctors at Diagnosing Illness
nytimes.com
nytimes.com
I was a big fan of the show House as a kid, and I remember being blown away when I learned that the “Department of Diagnostic Medicine” was made up for the show and not a standard department in every large hospital.
Replace AI with patient, and it's a far too familiar experience.
It's why the DSM in particular is so frustrating because the diagnostic criteria are written for how the condition presents from the outside. Which isn't inherently a bad thing but there's a huge amount of overlap which makes it imprecise. The criteria from the perspective of the sufferer is far far more specific.
> from the perspective of the sufferer is far far more specific
So yeah, that's... sometimes true.
Not a bad role for AI. Then again, you miss the chance to show empathy (this is me saying this as a radiologist who sees like 1 patient a day though so huge grain of salt)
``` Dr. Chen said he noticed that when he peered into the doctors’ chat logs, “they were treating it like a search engine for directed questions: ‘Is cirrhosis a risk factor for cancer? What are possible diagnoses for eye pain?’” “It was only a fraction of the doctors who realized they could literally copy-paste in the entire case history into the chatbot and just ask it to give a comprehensive answer to the entire question,” Dr. Chen added. “Only a fraction of doctors actually saw the surprisingly smart and comprehensive answers the chatbot was capable of producing.”
```
But I've also had medical professionals, particularly the non doctors (nurse practitioners, physicians assistants, etc) who are much less receptive and more fixated on their first guess, which has sometimes resulted in precious lost time and repeated visits for me. The linked research finding is interesting, and I think highlights the pitfall of professionals who believe too much in their own expertise or gut feeling even when they've not really examined the case carefully:
> The chatbot, from the company OpenAI, scored an average of 90 percent when diagnosing a medical condition from a case report and explaining its reasoning. Doctors randomly assigned to use the chatbot got an average score of 76 percent. Those randomly assigned not to use it had an average score of 74 percent.
> The study showed more than just the chatbot’s superior performance.
> It unveiled doctors’ sometimes unwavering belief in a diagnosis they made, even when a chatbot potentially suggests a better one.
> And the study illustrated that while doctors are being exposed to the tools of artificial intelligence for their work, few know how to exploit the abilities of chatbots. As a result, they failed to take advantage of A.I. systems’ ability to solve complex diagnostic problems and offer explanations for their diagnoses.
He was a history major before he went on to study medicine, and he now does a podcast on the history of medicine called Bedside Rounds. He gets really excited when talking about something he finds interesting and it makes you want to follow him down the rabbit hole. Highly recommend listening at half speed: http://bedside-rounds.org/
First: which chatbot can correctly pick the labs, imaging and other methods of investigation without wasting tons of $ and going off the rails with rabbit holes?
Second: get a chatbot to understand the clinical impression and correlate it to the history, labs and imaging.
Then: get a chatbot to understand that despite X being a standard antibiotic regimen for the infection, given the person's age, lab findings, and severity of the disease, Y for Z many days is actually a better strategy for these specific instances.
————————
Strange answer.
The patient has post procedure back pain and lower extremity pain (likely femoral access for his angioplasty) and acute kidney injury. Normally I would go for a few things such as an iatrogenic complication like aortic dissection from the angiocath.
His anemia, back pain, AKI, and malaise are also be worrisome for retroperitoneal hemorrhage, especially with recent anticoagulation with heparin and catheterization.
I might also think of heparin induced thrombocytopenia and get a HIT panel given recent heparinization and anemia. Although I’m pretty sure it’s not this, info was given in the stem for it.
The patient already had a CABG so he’s got multivessel coronary artery disease, now he’s also had a POBA. Does he also have ischemic cardiomyopathy? It wouldn’t be a surprise - acute decompensation could cause his AKI due to renal perfusion deficit.
And lastly to get back and extremity pain with renal failure after an angio sounds like the catheter would have needed to disrupt a plaque in the descending aorta somewhere and then shower the kidneys and lower extremities with emboli.
Anyways I’m a radiologist not a clinician so I’m sure there’s plenty of things I’m not thinking of. Plain old multifocal cholesterol embolism must be more common than I thought. So maybe strange for me but not for GPT or others.
Clearly not. These results show that most of the time doctors should be A.I extenders, offering valuable second opinions on diagnoses.
It tells you everything that he's "shocked", and despite that shock, he still maintains the above keeping with the cognitive dissonance. Many of us who have enough experience with modern healthcare could see this coming from miles away, and would have seen the opposite result (doctors beating GPT on average) as shocking.
50 sample size not very reliable.
The result that the human+AI is a little better but the AI alone is much better matches the experience with chess engines, where grandmaster+AI is a little better but just AI is the strongest.
To its detriment, “defeated” in the title is clickbait. It reminded me of those YouTube videos: “Liberal doctors DESTROYED by ChatGPT”.