Google AI has better bedside manner than doctors – and makes better diagnoses
nature.com
nature.com
Odds are you don't even see a GP anymore in the United States, you see a Nurse Practitioner who then potentially forwards your information to the doctor. Most visits your doctor spends less than 5 minutes on you.
This assumes you even have a GP already, since most are booked out 3 months or more for new patients. Reality is, the AI doctor you CAN visit is better than no visit at all.
https://www.niskanencenter.org/the-planning-of-u-s-physician...
Being a general practitioner is a hard job, requiring a long training. To keep good numbers, you have to invest continuously in the system and pay well. Unfortunately, this clashes with the ideological abandonment of the Welfare State post-Reagan/Thatcher/USSR fall.
Nurses in this case.
This is so true, and unfortunately can cause an NP to NP to NP circle of misdiagnosis
Yes, but it’s a problem if this becomes the goal. When the goal should be to allow everyone equitable access to the best healthcare.
I’m afraid that we’ll settle for subpar AI healthcare for the disadvantaged because “it’s better than nothing”.
> Yes, but it’s a problem if this becomes the goal. When the goal should be to allow everyone equitable access to the best healthcare.
> I’m afraid that we’ll settle for subpar AI healthcare for the disadvantaged because “it’s better than nothing”.
Yeah, and that's one big reason why "AI" will fail to live up to the utopian sci-fi hype that sustains its enthusiasm: our society lacks the ideological framework for those results. All "AI" will do is deliver more of the same.
Mark my words: AI will be the next offshoring: cutting costs (and jobs) by sacrificing quality and giving us (the plebs) no choice in the matter. Being consigned to live in a cardboard box under a bridge, consoled by an "AI" therapist like ELIZA, will be defined as success.
The hours are usually pretty constrained, the therapists are not great in general, they'll cancel last minute all the time. I had one show up clearly drunk. You'll get therapists that tell you they only do text messaging, no phone call or video chat.
Found a lady whose profile looked like it aligned pretty well with what I was looking for, but she apparently rejects all appointment requests from men (had my wife try to make one and the lady accepted it within minutes).
It's a shit show. I have no doubt AI therapy will become a thing.
I'm having a hard time figuring out your point. Is it 1) AI therapy will be good because your experiences have been so bad or 2) the ideological framework we currently have corrupts everything and has already corrupted therapy.
Honestly, after reading your comment and the Talkspace wiki page. It sounds to me like a garbage product in much the same way I expect "AI" to be: cheap through compromised quality, therefore favored by the powers-that-be.
Beyond price and ease of access, talking to a robot that isn't even capable of judging you might even be more appealing to some people.
Would a therapy tuned gpt-4 be adequate? Probably not, but who knows what the landscape looks like in a decade or two.
This is all obviously US centric and a biased opinion based on my personal experience.Talkspace is currently the best among these products in my experience, but that's only from a UX perspective, they all suffer from the same fundamental Uber for Therapy problems.
So basically, a repetition of the exact same point made upthread about nurse practitioners and AI? Instead of the "goal should be to allow everyone equitable access to the best healthcare," we choose "subpar AI healthcare for the disadvantaged [or merely non-wealthy] because 'it’s better than nothing'."
I should also note that therapy only came up in a joke mocking the myopia of too technology focused ideas of progress.
> Beyond price and ease of access, talking to a robot that isn't even capable of judging you might even be more appealing to some people.
I suppose, for a small subset of physiological problems and a certain uncommon types of people (which are likely vastly underrepresented among software engineers, especially those on HN), but my guess is that the impossibility of any kind of human connection is going to be a huge negative for most.
Regardless of the United States, this is the reality in poor countries right now and no amount of policy can fix that in the short term. I am excited for AI in medicine for this reason.
I don't like it any more than anyone else, but the reality is that the most likely thing to happen in the US is an explosion in the number of nurse practitioners and PA's.
Why? Because insurance companies will reimburse for them. Usually due to a myriad of reasons like standard of care, the insurance company having someone to lay hands on in the event things go off the rails, and a galaxy of other issues too byzantine to go into in a single post.
Until we wrest some control away from the payer, we'll see more and more NP's and PA's in the medium term. Fewer MD's. And AI's will struggle to gain acceptance as an arbiter.
In the far future, assuming the payers remain all powerful, you can see things going to a really natural looking place for the US, but a place that is dystopian in the extreme if you take a step out of the system and look at the big picture. Think of it this way, What happens in a far future where payers decide which AI's they reimburse for in the manner they currently decide which providers they reimburse for?
Replacing GPs with NPs will inevitably cause preventable deaths, the question is simply how many deaths we're willing to tolerate. That may be a perfectly reasonable tradeoff in health economics terms, but we have to be frank about it. I am hopeful that AI will go some way to bridging the gap between healthcare supply and demand, but at least in the short term we face a lot of difficult choices about how to allocate resources.
There's only one small thing: my appointment is on January 7, 2025. That is NOT a typo. All I have to do is stay alive another year and I'll get to meet my new doctor! Can't hardly wait!
Oh, I almost forgot: I'm a retired physician (neurosurgical anesthesiologist x 38 years).
PCPs are notorious for misdiagnoses, they're expensive, hasty, often times don't believe or listen to patients, and frequently just don't have all the data. This isn't necessarily their fault -- they're overworked and the medical industry isn't making things better -- but the reality is primariy care isn't working very well right now even in developed countries. Imagine in developing ones...
Diagnosing a patient is in many ways an expert system problem which computers are excellent at. Amassing the data from every medical textbook ever written, plus every study ever done, plus clinical conversations with patients and their medical history (the hardest part), and you have the best PCP ever made. Add a nurse to manage the physicality of it and one day connect data to the system to track lifestyle behaviors, sleep, and things people won't necessarily self-report, and you have something revolutionary. And AI is kinder and more compassionate, as the article said.
No wonder people have been trying to crack this nut for decades (albeit with minimal success). I hope the LLM revolution helps make another big round of progress.
Side note. The terminology seems a bit confusing. Wikipedia says "LLMs are artificial neural networks following a transformer architecture." It's a bit strange to call it LLM then and not "Transformer-LM", imho.
If you take a dense (fully connected) neural network and take away edges, you can end up at the transformer architecture. Perhaps IBM just used fully connected networks and an insane amount of computational power and used the transformers without even knowing it (?)
The point here being that adding AI diagnostics might improve on the quality of diagnosis but it might also potentially derail healthcare to some degree if the AI system doesn't question weather an investigation or treatment is actually worth it and should be prioritized. Then again, it might also be possible to make it prioritize more consistently and fairly...
I'm admittedly not a doctor, but this doesn't really match my understanding at all–after my dad got sick a couple of years ago I developed a bit of a fascination with reading about the practice of medicine, which has largely changed my view from an engineer's perspective like this to one with much more nuance.
Diagnoses in general are not nearly as cut and dry as people would like to believe, and getting to them is not often as simple as being a function of X symptom and Y test result. Patients are often vague or simply not equipped to provide a perfect history, tests have ranges and associated error, as well as risks of their own, treatments have risks themselves that may interplay with myriad other life factors. In many situations there may not be a definitive diagnosis to be had at all.
Imagine never needing to take another blood test again or getting early warnings for potential cancer that would have cost thousands to obtain before...
> It hasn’t been tested on people with real health problems — only on actors trained to portray people with medical conditions.
Obviously the FDA won’t allow them to operate in real patients, but designing a system that knows how to identify which script you had your actor play is very different from designing a doctor.
> an LLM has the unfair advantage of being able to quickly compose long and beautifully structured answers, Karthikesalingam says, allowing it to be consistently considerate without getting tired.
The last thing I want is to read through pages of AI blogspam at the doctors office. Where’s the AI that takes its pages of fluff and boils it down to the point?
That said, I am cautiously optimistic, as there’s definitely room for improvement. I’m lucky to be basically the same demographic as doctors, so they tend to actually listen to me and believe me. I know from second hand experience that for folks who don’t look like their doctors, getting one to take them seriously can be herculean. One individual I know had the UCLA (male) medical staff call security to escort her out of the building because they thought she was being too hysterical and should just tough out whatever it is she was dealing with (they refused to look). She then went to a (female) private practice gyno who took one look and immediately saw there was indeed a massively pain-causing complication present.
Let’s hope the AI doesn’t train on the wrong transcripts though…
Aside, the best part is UCLA didn’t want to let her graduate until she paid the medical bills from that “visit”. She ignored every one and they seemed to forget about it. Medical incompetence has its upsides.
Aside aside, this is also why “Shouldn’t the best test scoring individuals be admitted to (medical) schools, regardless of demographics? To hell with diversity, you want the smartest doctors possible, don’t you??” is flatly invalid.
I think you underestimate how much of an effect the doctor's style of response can have on patients. A curt response and an embellished response may both provide the same facts, but the latter is more likely to resonate with people. At the end of the day, everyone thinks they're special and not just "another patient".
I am not advocating for "pages of AI blogspam" -- LLMs can be instructed to craft better responses than that -- in the real world, doctors suffer from empathy "drain" very often, LLMs don't.
That's all well and good if the patient doesn't know they're interacting with an AI though!
To clarify, my stance on this whole thing is mixed.
I'd much rather see something like this being owned by a government of some sort, building it in the open and without corporate incentives.
Relative to the average patient, you are an extreme outlier. You are almost certainly well above the population average in terms of your level of education and cognitive ability. The average patient is well below that population average, because healthcare is very disproportionately needed by people who are older, less educated, have English as a second language, are suffering from cognitive impairments etc etc.
The number of people who complain that their doctor spent too long talking to them and gave too thorough an explanation rounds to zero, which makes an infinitely patient and free-at-the-margin AI doctor an obvious improvement. A well-trained LLM will make the same sort of assumption about you that a human doctor does - this guy presents his history concisely and in clinical vocabulary, so I can probably skip the pleasantries. The LLM should (if trained towards the right outcomes) not make the sort of prejudiced judgements that caused your friend's unfortunate experience. If it does, we'll see it in the data.
All the while ignoring that the medical field refers to "misdiagnosis leading to death" as "Wednesday".
https://www.propublica.org/article/cigna-pxdx-medical-health...
> The company has built a system that allows its doctors to instantly reject a claim on medical grounds without opening the patient file, leaving people with unexpected bills, according to corporate documents and interviews with former Cigna officials. Over a period of two months last year, Cigna doctors denied over 300,000 requests for payments using this method, spending an average of 1.2 seconds on each case, the documents show.
Bull. One of the challenges of the medical profession in the XXI century is that google-assisted patients are ready to sue at the minimum sight of anything even slightly off the perfect diagnosis and treatment.
Yes, a lot of doctors are assholes, often with a god-complex - not unlike every other human being out there - but the fact that you can sue the assholes keeps them increasingly in-check. AIs have no conscience and can't be sued - you can sue the parent company, at which point it will just become a business cost, in the same way chemical certain companies put aside a bit of cash to pay before they discharge crap into rivers.
No lawyer will take that case unless you pay cash up front; if you did, they'd be professionally obliged to tell you that you're wasting your money.
Medical malpractice is defined and decided wholly in terms of the accepted standard of care. If you did what your competent peers would do in the same situation, if you followed the guidance of a recognised expert body, and if you documented that thoroughly, then you're legally in the clear. A great deal of medical care is demonstrably sub-optimal, but still very comfortably above the threshold of negligence.
Overwhelmingly, patients don't sue because they have over-inflated expectations - they sue because clinicians and hospital systems make a lot of foreseeable and consequential errors.
We've got ardent supporters and detractors, and often with this type of divide reality is somewhere in the middle. This is a huge, dangerous, and sometimes hard to understand technological development.
This is a hand-wavy doomer argument. This is why we have trials and statistics - if constructed properly those should produce a quality answer to the question of using such systems. If on average it gives the same or better results than an average human doctor then it is good to go.
...but like self-driving cars the errors can be terrible even if it's only 0.1% of the time. I might not be as vigilant as AI-assisted driving, but I'm also not going to get confused by a truck carrying a stop sign and slam on the brakes on a highway. On the macro scale improving the average is great, but on the personal scale I can't only trust the average.
An LLM without question can be better than a human much of the time, but the errors, while more rare, can be worse than human error due to the lack of contextual reasoning and general intelligence.
That is simply personal bias and has little to do with reality. The basis of e.g. the scientific method is rejecting what you "trust" or not when having good quality statistical information which suggest something else.
> An LLM without question can be better than a human much of the time, but the errors, while more rare, can be worse than human error due to the lack of contextual reasoning and general intelligence.
Do you maybe have anything to support this claim or is this simply your personal feeling/belief?
Anecdotally I use LLMs every single day, and almost every single day it makes a silly error that many humans would not make because it's not continually reasoning or interacting with reality. Until I see these silly mistakes go away, I will always work alongside it to verify output. I'm not giving it the wheel of my car anytime soon.
maybe if you are rich. For most people medicine is going to a GP once or twice a year, spending an hour in a queue, then seeing a doctor for 10 minutes.
They want me to talk as little as possible, and I want them to fix whatever problem I have (or look for problems I can't detect, like high cholesterol, etc). It's utterly transactional and would be greatly improved by an AI that I can work with on my own time to build a real medical history, not a 1-page sheet that I need to fill out four times a year.
I assume the major challenges that will be difficult to solve will be similar to what they're already facing, namely dealing with patients who can't communicate their issues clearly (or correctly) or who are being deliberately misleading e.g. with drug-seeking behaviors.
In general looks like diagnosis will be in large part overtaken by ML, the amount of knowledge you can cram into it is orders of magnitude more than with the smartest humans. And doctors (especially good ones) are very expensive, with this even some mediocre one can produce amazing results.
“the base LLM with existing real-world data sets, such as electronic health records and transcribed medical conversations”
So it trained on text that a healthcare provider extracted from the HPI/interview, labs/tests, radiology etc?
So it’s just a diagnostic assistant that can chat? Is it going to get conversational info out of non-English speakers and indigents? The urgently ill who are unable to vocalize what they’re feeling? The delirious who have “altered mental status”?
What about the huge amount of geriatrics who have significant cognitive decline and lack the ability and vocabulary to deliver information? And this is going to bother the terminally online - but what about the low SES who speak in virtually indecipherable urban patois?
This piece of shoftware isn’t going to help anyone. It’s going to be for articulate young to middle age educated people who don’t want to talk to a chatbot. And they’ll probably hate using it as they chat with it for filling out visit screening and new patient forms. If it ever deploys.
An “underserved” patient may use it out of desperation to get what they need but I would never in a billion years want this for a patient I gave a shit about.
I guess there’s the crux. They don’t give a shit. Some PE group will get it and you can talk to it before your virtual visit with your human doctor on Zoom. It’ll be like Elysium. Maybe I’m bitter because I still commute all over town. Maybe that’s what all you software people want, you like WFH so much, let the doctors do WFH too eh? Talk to the chatbot, you guys made it and it’s better than us medical cartel types anyways.
> An artificial intelligence (AI) system trained to conduct medical interviews matched, or even surpassed, human doctors’ performance at conversing with simulated patients and listing possible diagnoses on the basis of the patients’ medical history1.
> The chatbot, which is based on a large language model (LLM) developed by Google, was more accurate than board-certified primary-care physicians in diagnosing respiratory and cardiovascular conditions, among others. Compared with human doctors, it managed to acquire a similar amount of information during medical interviews and ranked higher on empathy.
So, they created an interview machine? Getting better than a medical doctor at conducting a dialogue based interview is simplistic and sophmoric.
Primary Care docs are the most generalist doctors that exist, and medicine is very much a specialist field.
2nd, doctors do not rely on only dialogue to evaluate a PT, but instead they use in situ observations and sensing devices.
On the good side, this could probably help replace some Nurse Practicioners working for insurance companies that have no business getting anywhere near a PT.
That's technologists for you: redefine the problem to suit the strengths of your technology, then paper over that in your hype.
That being said, I had to laugh, this is one hyperbolic headline. There are of course some caveats.
"Few efforts to harness LLMs for medicine have explored whether the systems can emulate a physician’s ability to take a person’s medical history and use it to arrive at a diagnosis. Medical students spend a lot of time training to do just that, says Rodman. “It’s one of the most important and difficult skills to inculcate in physicians.”
This does take a lot of training and experience. The challenge in diagnosis isn't about asking a history and integrating the information, it's about effectively encouraging patients to provide necessary information that they might not realize is relevant or know how to articulate. Different patients often require wildly different approaches. Medical literacy does play a role, but a patient that would say, "Hi doctor, I experienced central chest pain accompanied by discomfort in the upper stomach that happened two hours ago" (from pg 33 of the preprint) is not realistic. More likely you get a vague complaint of "heartburn that started a while ago".
Similarly: "Currently, I'm not on any prescribed medications". More frequently you get something like "I take a blue one" or "Golly Telly" (Go-Lytely) or "Gabatini" (Gabapentin). I do think an LLM could probably parse these but such idiosyncrasies in a history do compound. And though someone may be prescribed a med and think they take it as prescribed, sometimes it takes a hunch and clinical experience to tease out that in fact the med is not being taken at all as indicated.
Moreover, the better bedside manner was assessed via text conversation. I wouldn't quite call that "bedside" manner. I also wonder how such a system will deal with patients that have self-diagnosed themselves and looked up the "right answers" to get what they want -- a difficult reality that takes some parsing to figure out what's real and what's not.
Overall though, the preprint, in contrast to the Nature news article/headline, does a better job of discussing these limitations. I congratulate them on excellent work. Thank god LLMs can't do surgery, though I'm sure my time will come as well.
This study is so far below the standard of evidence for medicine it's comical.
1) Take a given ailment
2) Take all the symptoms that it produces
3) Choose some at random
4) Construct a synthetic patient case with those symptoms where you know exactly what the problem is
If that was really an issue for you, you'd have gone without medical care. Anyway, you consented by electing for medical care, so it's all fine, so very fine. /s
What? What kind of "bias" would this entail? Or is this more of the "ethics" garbage that is used to lobotomize LLMs? The real scientific-technical problem and discussion about aligment (e.g. paperclip maker) has been replaced by some strange grift that opens doors for people without any useful knowledge into the New Big Thing where the money is.
Something like half of medical students still falsely believe black people feel less pain than white people: https://www.nytimes.com/interactive/2019/08/14/magazine/raci...
Some sensors don't measure accurately on darker skin: https://www.chop.edu/news/oxygen-readings-may-be-affected-da...
etc. etc. etc.
If I construct a ML system to classify animals into kingdoms but my data is lacking e.g. cats so it skews the results away from the goal is that an ethical issue?
You must correct the bad data before using it for actual patient care to avoid an ethics issue. This means we must consider the potential ethics issues of its use prior to implementation, right?
> If I construct a ML system to classify animals into kingdoms but my data is lacking e.g. cats so it skews the results away from the goal is that an ethical issue?
If you wind up mistreating cats as a result, sure.
Is this better than no healthcare at all, or having to wait a month for a GP appointment? Likely yes. Is this good enough in terms of what healthcare should look like in a first-world country? Nope, and never will be, not even with healthGPT v77.