Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine
nature.com
nature.com
I'm really looking forward to studies that look at performance comparisons in realistic environments. I believe there is a potential revolution brewing for PCPs. Having someone who actually listens and has real in-depth knowledge that goes beyond normal med students could be a game changer when compared to the usual conveyor belt care that gives everyone the same diagnosis based on a quick glance.
... or if you've got anything that requires know-how/practical experience, which coincidentally is what makes most of medicine.
That makes a world of difference. It's actually where most of a doc's performance resides.
> in a clinical setting you won't have the expert chief physician re-do the physical either
And you don't need it. Lots of thousands of fellows or even some advanced resident are already at the approriate level of experience. Actually, fellows are often superior because they've got a superior volume of practice.
Got any source on that? In the paper they certainly don't assume that.
>Lots of thousands of fellows or even some advanced resident
Same argument as before. We simply don't have enough of those in western medical care due to aging societies.
I would argue the opposite.
Any reasonably well trained AI should outperform doctors on rare/exotic things, that most people will see never or only every other year. As it will still be described in the literature, the AI will have picked it up, but our average doctor is more likely than not to have forgotten about it since med-school, if she ever heard about it in the first place.
But when it comes to the normal day-2-day stuff, I am more afraid of the AI having a sudden case of hallucinations vs the doctor who probably has seen that same case yesterday and might not apply the latest state-of-the-art treatment [1], but it will probably be fine.
[1] It often takes 5 years or more for the medical state-of-the-art to transpire into daily practice "in the field".
Edit: Add citation and fix typo.
This is actually super dangerous and probably the cause for a lot of misdiagnosing and medical errors these days. From the paper, it appears that GPT does much better in these cases.
[1] Sure, you have to test your new cancer drug on actual cancer patients, but why include the 65 year old obese patient, when you could instead include the 32 year old triathlon athlete into you cohort?
It's not just better than a blind man. It's better than any human, once you scale the task up beyond a certain point. Consequently, we should not get hung up on the path taken if the ultimate goal is the best possible care for patients in our overloaded public systems. And this paper suggests there is great potential.
Although I would tend to say we should expect machines to be significantly better than the average doctor. Before we role them out to everyone we should require them to be at least as good as the top 10% of doctors, I guess.
But regarding diagnosis, there are symptom checkers where you feed stuff in. I get the impression that multiple companies are in the process of building such tools based on crawling research papers to diagnose even very rare conditions. IIRC IBM Watson had such functionality.
I mean, it's not like rocket science. More like investing the effort to build a really big decision tree.
If there is some evidence they are gathering that ChatGPT doesn't have access to, an effort should be made to capture and systematise it. But there probably isn't very much of that for a typical doctors visit.
You don't want your doc to give your problem his/her best shot !?
> used a checklist and evidence based process that can be engineered and improved over time.
Enticing and logically evident, isn't it? Well, except in practice it would cost so much to do that, that it's clearly not economically doable. And I'm not even sure it's technically doable in 2024.
> If there is some evidence they are gathering that ChatGPT doesn't have access to, an effort should be made to capture and systematise it. But there probably isn't very much of that for a typical doctors visit.
GPT is missing most of the info, because it's not equipped with 5 human senses and a brain to aggregate them. And that's, by far the most important info. You can see the challenge in solving this problem at a large scale, I'm sure.
I don't think the doctors gut feel is going to be their best shot. I want them using logic, evidence and reason. The medical industry has a long history of people slowly managing to convince doctors to use evidence instead of making bad guesses. Explicit processes generally lead to far superior outcomes.
You're not wrong about what the present is, but you are misrepresenting what a desirable future looks like. We want to engineer the guesswork out of the system altogether.
How would I know beforehand if I have something fit for an AI or for a human?
Are you saying just always assume you have a common, non-peculiar issue, and if I happen to have the uncommon one, bad luck, got the wrong treatment? At least take solace in the fact that it doesn't happen often?
This is what most discussions seem to disregard. Sure, a celebrated expert on a good day with prep time may still outperform a machine learning system. They will also outperform a mediocre practitioner at the end of a long week, probably by a larger margin.
When I mean bad, I mean that it cannot identify properly if the bone is human or animal, or which bone it is, and describes something completely different.
The benchmarks are too optimistic in such cases.
{image}
{question}
{choices}
Please first describe the image in a
section named “Image comprehension”.
Then, recall relevant medical knowledge
that is useful for answering the question
but is not explicitly mentioned in a
section named “Recall of medical
knowledge”.
Finally, based on the first two sections,
provide your step-by-step reasoning and
answer the question in a section named
“Step-by-step reasoning”.
Please be conciseI am dealing with serious post COVID complications (Long COVID). Doctors in most places don't seem to be taking this condition seriously and believe it to be a made up thing, but COVID causing damage to bodies and brains is a very real thing.
Thus, for a lot of long COVID patients, treatment has to largely be figured out by oneself, often through reading upcoming studies and experimental treatment reports from online forums.
ChatGPT has been really helpful to me and other patients I know to help make sense of and get context around medication, supplements or other treatments we hear about. I'm not usually asking it what to do, but it really helps me understand what a given medication (say sulbutiamine or whatever) does and how it works and what potential interactions it has with other medicines etc.
The other day it helped me figure out which amino acid complex to take, I put in the nutritional information of different products and had a long conversation with it to figure out which one is a better fit for me, and then confirmed this with external web searches.
I've also used it to help figure out what bloodwork to get and what diets or exercise routines are feasible for someone in my position (with limited mobility and strength and other resources). I can confidently say that I would be much in a much worse place than where I currently am with my health if I didn't have ChatGPT to consult with.
It gives me important context and information to help me make informed decisions about my health.
These are conversations I'd ideally have with a doctor but at least ChatGPT fucking listens and knows some stuff. I don't want to have to rely on a damn chatbot for my health, but high quality doctors aren't accessible to me due to inequalities of this world. So while y'all work on that, I'mma do what I can to survive.
What happened with "trust the experts"?
It was always lazy shorthand for “we don’t fully understand what’s going on but in general you’re probably better off by trusting the current scientific consensus,” which is a consistently good approach across time and space even if the consensus turns out to be wrong.
And for cases where there isn’t broad consensus, like the etiology of Long COVID, then obviously “trust the experts” is an almost semantically meaningless statement anyway. This is not analogous to climate change or vaccines or other areas where you hear this (again, lazy) phrase a lot.
The results were obscenely strong.
So no, not really. RCTs of that scale and with those strong of results build consensus pretty quickly, which is why there was (and remains) such strong consensus, which is why you feel like "trust the experts" was such a cudgel.
These arguments are always fun because they resolve to, "I'm mad because I wasn't trusted with subtle arguments and statistical nuance" immediately followed by a demonstrated inability to grasp exactly those things. Of course if you really wanted the subtle arguments and statistical nuance, you could've looked at the study results themselves. But you almost certainly didn't and most people wouldn't, ergo "trust the experts" who do those things is often very reasonable advice.
I am just personally wary of diagnosing myself - what I may think of as Long Covid (because it correlates to symptoms I've read about) may just as well be something else. I have no way of knowing.
But everyone's circumstances are individual, I suppose. Especially when regular doctors are unable to help you. That's a tricky situation.
Gemini and Llama have no such customer noncompete.
Training new models on ChatGPT conversations was a pretty common way to "cheese" getting local models up to snuff (or was, before they got there on their own)
I can't imagine it somehow binds you to never redistribute the knowledge you now have in your head, thats literally the same argument these AI companies have been making as to why its not copyright infringement when their model scoops up the IP of everything on the internet and regurgitates the amalgamation.
The idea of people using them for medicine is frankly terrifying to me.
This reminds me of the self driving car debate.
Plenty of medical people are wrong in hard-to-spot ways. Plenty are wrong in easy to spot ways too.
If we get a machine to do the job, should it be perfect, or should it be better than the average human?
I'm not expecting perfect, I'm expecting not literally every non-trivial response to be subtly wrong, which is my experience with code from them. If every time your self-driving car goes out it runs someone over, no, we shouldn't allow it.
GPT-4 is often scarily good at suggesting diagnostic and treatment plans for even very rare conditions. It will sometimes suggest things even well-trained doctors wouldn’t think of.
But it also makes some basic and expensive errors. For example, it often suggests long COVID patients with tachycardia get an echocardiogram (which is unnecessary and can cost hundreds of dollars).
So I think the model where AI drafts, but a doctor edits, would lead to much better patient care.
Only thing that I found that helped was a strict carnivore diet and weight lifting (med-heavy) for a month and all of the symptoms subsided. I put more stock in the weight lifting, YMMV.
The practice of medicine does not produce cute little cue-card prompts with 4 options. Diagnosis means trading the risk of various interventions against their diagnostic value, including a wait-and-see strategy. It means consulting with the patient on impact, as well as conferring with other specialists.
It's of less-than-zero value to have an 80% machine accuracy and 75% clinician accuracy if the impact of the machine's mistakes are high risk, unconcerned with the impact of the intervention on the patient, or provide little therauptic upside. Likewise, on these decision benchmarks *Machine decisions are NEVER iterated* -- this makes any claim to "performance" not even borderline pseudoscience, but imv, malpractice.
In the realworld practitioner decisions are iterated: doctors do not make high-risk mistake on every sequential decision, given more diagnostic information. It is highly highly unlikely that an inaccurate high-risk judgement call is compounded on each intervention. Whereas machine decisions are never tested in this sequential manner -- if they were, their benchmark performance would drop off a cliff.
The production of non-iterated decision benchmarks which measure accuracy and not risk-adjusted real-world impact constitutes, imv, basic malpractice in the ML community. This should be urgently called out before non-experts think that LLMs can give credible medical advice.
Accuracy and other quantitative performance metrics are imperfect for sure, but how else do you propose testing before real-world deployment? How do you propose scalable and feasible testing of human students?
> The practice of medicine does not produce cute little cue-card prompts with 4 options.
Ah, but the New England Journal of Medicine's Image Challenges (designed to test the knowledge and diagnostic capabilities of medical professionals) does.
> It's of less-than-zero value to have an 80% machine accuracy and 75% clinician accuracy if the impact of the machine's mistakes are high risk, unconcerned with the impact of the intervention on the patient, or provide little therauptic upside
But this paper does not study the practice of medicine. It intentionally focuses on performance on one specific, well-known medical imaging diagnostic challenge.
95% of ML "research" would reveal itself pretty useless if people did this, since the "80-90%" accuracy we're getting is on the "break-even" part of the profit-curve. We're not getting innovation.
It's very rare that frequentist stat modelling on historical data would produce anything, in this sense, suprising -- ie., suprise in the utility domain
It’s not to say the GPTs are currently useless though, they can provide a starting off point or remind us of things we may have missed. Full iterative diagnosis seems really far off though.