63 karma · joined June 22, 2021
That being said, I had to laugh, this is one hyperbolic headline. There are of course some caveats.
"Few efforts to harness LLMs for medicine have explored whether the systems can emulate a physician’s ability to take a person’s medical history and use it to arrive at a diagnosis. Medical students spend a lot of time training to do just that, says Rodman. “It’s one of the most important and difficult skills to inculcate in physicians.”
This does take a lot of training and experience. The challenge in diagnosis isn't about asking a history and integrating the information, it's about effectively encouraging patients to provide necessary information that they might not realize is relevant or know how to articulate. Different patients often require wildly different approaches. Medical literacy does play a role, but a patient that would say, "Hi doctor, I experienced central chest pain accompanied by discomfort in the upper stomach that happened two hours ago" (from pg 33 of the preprint) is not realistic. More likely you get a vague complaint of "heartburn that started a while ago".
Similarly: "Currently, I'm not on any prescribed medications". More frequently you get something like "I take a blue one" or "Golly Telly" (Go-Lytely) or "Gabatini" (Gabapentin). I do think an LLM could probably parse these but such idiosyncrasies in a history do compound. And though someone may be prescribed a med and think they take it as prescribed, sometimes it takes a hunch and clinical experience to tease out that in fact the med is not being taken at all as indicated.
Moreover, the better bedside manner was assessed via text conversation. I wouldn't quite call that "bedside" manner. I also wonder how such a system will deal with patients that have self-diagnosed themselves and looked up the "right answers" to get what they want -- a difficult reality that takes some parsing to figure out what's real and what's not.
Overall though, the preprint, in contrast to the Nature news article/headline, does a better job of discussing these limitations. I congratulate them on excellent work. Thank god LLMs can't do surgery, though I'm sure my time will come as well.
I value the former and find ways to discount the latter. So I am very happy. Though sane or not would be up to others.
With 60-70% correct rates on most training sets and 0.63 critical errors per report, for any physician not very well-versed with the limitations of LLMs, this is more of a liability than an asset. Some of the biggest barriers to care are cognitive, such as anchoring or availability biases. LLMs in their current state will only muddy the water.
Good physicians already know and do use these tools, bad ones will only get worse. A legal mandate will not benefit care.
Doubtless these models will progress to where this calculus will change. The only benefit from a mandate now that I can foresee is to accelerate fine-tuning by forcing widespread reinforcement learning by physicians, but that is a different discussion.
A common analogy used in neurosurgery training: "Imagine a wooden plank on the driveway and walk its length -- no problem. Suspend that same board ten stories in the air and try again." The hard part is not the "figuring out" or "execution" but knowing the irreversibility and making the correct decision. Your patient expects you to be correct 10 times out of 10, yet you know that's not possible. Squaring our fallibility with the irreversibility of our missteps is the hard part, and what keeps surgeons up at night. The fly struggles in the web.
I'm a neurosurgeon and we commonly advise our patients: Even after a "perfect" or minimal surgery without any evidence of periprocedural stroke or complication, you may not ever be the same again. Sometimes it takes months for these changes to be noted, sometimes it's only even noticed by family members. Odd word finding difficulties, perception changes, memory/concentration issues; the gamut is endless. As we say, no one's the same when the air hits your brain.