Hurts who ? Yes the doctor is super stressed and has maybe 10 minutes for you that's the actual problem, it's not like before LLMs they were super glad to sit there and answer all your questions.
I wouldn't trust AI to make a diagnosis, but I would absolutely trust it to notice where procedure hasn't been correctly followed, where a treatment is counter-indicated because someone has missed a line on a health record, or where there's a clear potential alternate diagnosis which has been missed for spurious reasons. Also, unfortunately, where doctors aren't doing a decent job - often because they're overworked or underfunded.
The same issues that were present with search-engine self diagnosis are still present with LLMs. If you provide Google with an incomplete list of symptoms and can’t interpret the information you find correctly, you will likely get an incorrect diagnosis. The same is true for LLM output.
There's a reason I ask AI about absolutely everything medical and there's a reason I keep extra quantities of prescription medications around for emergencies. I've saved my own ass a lot more times than the doctors have, thanks to good doctors not being available.
I get it. But the current system is also super difficult for the patient: getting time to ask questions, get clear answers, get the best possible diagnosis taking into account your history, symptoms etc and all that in 5-10 minute checkup when your doctor sees 50 patients a day and has very little time for you; this doesn't scale well. Patients run to A.I for a reason.
But AI's problem is that its completely full of shit, sometimes, and the people most qualified to evaluate whether its full of shit are the doctors, not the patients, but just like OP's original article, patients are left feeling like their second opinion from AI might be more trustworthy than their doctors opinion.
Examples of things normal people can verify
- procedural errors that Claude can capture like some blatantly high dosage (grams instead of milligrams)
- outdated treatment plan, maybe there’s a credible new treatment plan that’s been used for years but the doctors were not updated
- literally being injected homeopathic drugs (takes no smart person to flag this)
Let’s stop talking as if doctors have a divine right here. And let’s accept some agency.
A doctor might have never recommended upping X, because they would know what it does to your body. Or they might have suggested additional supplementation to avoid this.
The fact that LLMs are trained on all public knowledge is a huge red flag, because there are more wrong infos out there than right ones. Especially about health, diet, etc.
It's now quite unusual that it's "Completely full of shit". If it contradicts something your doctor said I don't see why you should feel ashamed to bring it up. Sure it complicates the doctor's work, having ignorant obedient patients must be more comfortable for the doctor, but the end result could be more accurate diagnosis.
Studies have found that newer reasoning AIs are about as good at diagnosing illness from a written description of symptoms as doctors are.
Granted, it cannot actually examine a patient, so we're not replacing doctors anytime soon. But your view is obsolete.
It may have some utility after diagnosis, but this test doesn’t demonstrate utility for patients.
The more training data, the more questions it can answer with a reasonable degree of probability of accuracy.
Throwing away a potentially useful analysis just because it’s probabilistic seems a bit like throwing the baby out with the bath water.
This case is about handing a 3D imaging result to a text predictor and hoping for a valid second opinion.
The real question is where’s the cut-off point between accuracy and utility.
Remember: a second human opinion can also be wrong, and even a wrong opinion can still be useful (especially in medicine where differential diagnoses are a common practice - if the LLM gives you a useless opinion, you rule it out and move on).
I don’t think it’s particularly unreasonable to think that an LLM would have enough literature, or enough reasoning ability, to be able to generate a plausible interpretation of the data. A human can then review and say either “yeah that’s clearly not the case here” or “hmm, actually that could explain it, maybe we should order another test”.