Towards accurate differential diagnosis with large language models
arxiv.org
arxiv.org
I have a friend with Crohn's who was feeling low energy. I was a gym bro at the time and convinced him to take a testosterone test (because all problems re caused by low T when you're a gym bro).
His doctor wouldn't even entertain the idea, saying he's a young man and it's very unlikely that he'd have low T. He did the test privately and his T is significantly below normal. If you Google, there are actually many papers showing correlation between Crohn's and low T. I bet an AI would find it.
Similarly, doctors missed my mum's recent cancer diagnosis. She also had factors that would make her more susceptible to breast cancer, googling finds many papers that show causation.
The problem is that those things aren't extremely common and haven't made their way to NICE guidelines or whatever GPs use.
Not that I'm blaming doctors, they have 10 minute appointments and don't have time to do anything. I'm sure AI would recommend significantly more lab tests which would put even more pressure on the NHS.
So, I talked to my endocrinologist and he said it wasn't a normal test they do, but he was happy to do it if I wanted it. Well, the test came back positive. So they sent me in for a CAT scan of my adrenal glands to see if anything was obviously abnormal there. Which came back negative. So, now they want me to do a salt test to see how bad my hyperaldosteronism is, so that they can give me the right amount of the drug for that condition.
But the key thing that set me off was that original article that said this was now recommended by cardiologists as a standard screening test. Maybe only 10% of patients have this problem, but it's common enough that it's still worthwhile screening for.
So, why didn't my cardiologist screen me for it in the first place? Why didn't my endocrinologist screen me for it without me having to explicitly request it?
I'm pissed.
I don't want to sue anyone for malpractice, or even mention that word to any of my doctors. But I do want to convince them that they really do need to screen for this condition.
Meanwhile, I'm going to continue to follow this diagnosis to see where it goes, and if they can help me get my blood pressure under better control.
The only way to convince them is to sue. But there's nothing worth suing here for, you got the test. You can only sue if damage was done, then the doctor would be responsible for blocking the test when you explicitly requested it.
If you have any suggestions in the Austin area, please let me know.
Hopefully technology like this can help us rethink those decision trees, or else crunch more patient-specific information to adjust the diagnosis process on an individual level more efficiently.
Said to physicians: "I dropped out and I TELL EVERYBODY I KNOW you cannot pay a doctor enough for all the unnecessary sacrifices they had to make, just to prescribe you an antibiotic."
...to Non-physicians: "While I believe that you cannot pay physicians enough for their sacrifices, you should definitely investigate your own ailments, and seek multiple opinions whenever your `gut feeling` indicates `abnormal`. Physicians simply DO NOT HAVE TIME to give a fuck about your specifics.
More recently, I have included (to all) "it is probably ALREADY UNETHICAL TO NOT BE CONSULTING WITH LLMs during differential diagnoses."
Just my ¢¢
I've seen way too many people who I consider generally intelligent, completely falling for the convincing bullshit. Mostly in my online communities, but also a few times in person.
The problem is that respecting every detail of every guideline all the time is pretty much impossible, and even a debatable stance. There are many committees producing guidelines left and right. Because those people's jobs are to produce guidelines, not care for people. The exact same goes for research papers, and validity for patients in practice is a difficult issue.
So, docs all practice mostly from habit and experience. It's a craft supported by science, but clinical medicine itself is science as much as bricklaying is a science. But that doesn't make AI the use-all-do-all, because if everybody starts to test for every little thing, diagnostic probabilities will decrease massively and guidelines would have to adjusted in conséquence. It's a feedback loop. And also, you'd pay 5x the price for healthcare, so not realistic.
It's hard for me to write out all my thoughts on the subject, but I completely agree with this. The paper cherry-picked difficult cases and compared how AI deals with them vs regular clinicians. I think that is something AI will excel at because there is no cost to being "wrong", as long as something it spews out is correct.
In difficult to diagnose cases, that kind of makes sense. Find the answer at "any cost". But you cannot apply that to every case. If you let the AI loose on fairly simple cases, I bet it would over complicate many simple things, because it couldn't handle the nuance of people coming in and just whining about a regular cold. "Patient comes in with a headache, might be brain cancer but might be a cold" ~ type DDx.
I have already nearly stopped using Google search for anything, in favor of GPT-4.
GPT-4 has helped me very quickly prototype things that I normally would have had to spend hours researching.
GPT-4 has also created custom curriculum for me to help me learn various things for which I have struggled to find good books/tutorials online.
There will always be many areas in which a solid human intellect and well-honed human judgment is still useful, but much of the less critical work will yield to LLMs.
If we call the difference between GPT-3 and GPT4 1x, then I would expect to see a 2x-5x improvement in LLM capability within the next few years just based on how much great work has recently gone into shrinking really big models so they can run on smaller hardware.
LLMs are not digital human minds, they are simply very good information synthesizers. Information synthesis happens to be what most white collar professionals get paid to do with their brains.
While I don't know you personally and I'm not talking about you specifically, "just" feeling very excited about something doesn't Make This About You, the prognostications don't actually help you Be a Part of This.
This is the same energy as being really into COVID-19 (what was it called, "corona-scrolling?"), the gambling-fueled crypto boom, the retail stock trading booms. It's kind of adjacent to the fallacy Feelings that are More Strongly Held Are More Valid.
It's like you could use these things as a barometer for the health of a social media channel.
Anyway, I'm sure there's a name for this part of the hype cycle - tapping into how people want to be a part of the number one biggest, hottest trend by, at the very least, talking about it all the time, trying to compete against each other by sucking the most air out of all rooms through ever-greater hyperbole.
One thing's for sure: there's some guy who found some massive low hanging fruit in some neural network designs by carefully choosing brackets in chained matrix multiplications. That's one guy, and he definitely Made It About Him in a way that makes sense. How many people on Earth have enough knowledge to see things like that, maybe 1,000-10,000? It's so exclusive, in a sense.
I guess my point is, you're pretty far off the mark. In the interest of curiosity: Standard practice in medicine is to examine the patient before giving an opinion, a big roadblock for even the most competent of multi-modal models. Also, most people are biased in that they are intimately familiar with their own lived health problems, so they fill in tons of blanks, discounting the value of inquiry during a consult, whereas doctors may see you for less than a few hours a year, LLMs much less so.
At a minimum, I’ve already begun using gpt-4 as a consultant for my fathers medical issues. I can reasonably get confirmation/dis confirmation of whether a medical experts opinion is BS or not. As anyone who has dealt with the medical system in the US can attest - the majority of opinions you get from experts are either BS, outdated, or folks talking their own expertise.
Asking the question to GPT-4 can help you decide when to talk to a different doctor, or try something new entirely.
The last paragraph caught my attention though you were a bit abstract _and_ obtuse, due to the structure being "brief aside for the curious: LLMs don't have bodies"
My job and existence is centered around enabling doctors via LLMs and it absolutely is astounding to and for them. Funnily enough, it's the simplest thing: citations and easy UI to reach them. They perceive it as quick literature review on demand, as far as I can tell, that means they get to skip about 5 minutes across 4-8 articles checking links for relevancy.
I don't think OP was proposing that LLMs will replace doctors as you seem to think. I can't identify anything in their post suggesting that.
And if a few years down the line it turns out that you were wrong, everyone else was to, so no big deal.
In fact, you can already see it happening. Scaling up LLMs was surely enough to get to AGI, but now that apparently OpenAI is working on Q*, it's obvious that LLMs alone aren't enough, but LLMs + Q* are!
Will a good LLM based "consultation" fairly often have to refer you to a face to face consultation? Probably, but of my consultations over the last few years, a significant majority were resolved purely remotely.
it's interesting this is probably the assumption of many people leading to the current hype cycle.
But in reality it's not how progress tends to happen which is a much more punctuated equilibrium type phenomenon. My guess is it's actually quite likely that most of the progress will be made on maturing things and underlying capability will plateau.
In order to automate most medical care we'll need multiple breakthroughs that go way beyond LLMs. Linear improvements in LLMs won't get us there.
LLMs do have potential to improve other areas of clinical practice such as charting.
and for a follow up visit, read the prior visit's note and orders, search/track down all the test results (may be at more than one institution with multiple EHR systems that don't interchange) and summarize
or even just be able to take verbal orders like it used to be: "get a CBC, CMP, TSH, and CXR"
The NEJM CPC's are after the hard work has already been done to get the data to present the case
LLMs were born to do this.
> or even just be able to take verbal orders like it used to be: "get a CBC, CMP, TSH, and CXR"
GPT-4 of course knew what that meant, although I didn't. I told it to roleplay a potential follow-up:
> "Patient with fatigue, weight loss, mild fever. CBC shows anemia, elevated WBC. CMP reveals elevated liver enzymes. TSH normal. CXR clear. Consider ESR, CRP, ANA, and abdominal ultrasound. Possible infection or autoimmune condition. Referral to a hematologist may be warranted."
it's the taking of data from disparate sources including literal paper faxes that's the hard part
Making data ingestible is just laborious and unappetising to an engineer. So hand it over to a data entry operator who’ll do this happily for a wage.
I’m not sure why the issue of making data available to an LLM is a failure of the LLM.
Btw, OpenAI’s LLM can very well parse images, read handwriting, fix sloppy spellings, and extract objects from unstructured data.
Other than that dealing with disparate data formats is something LLMs can already do very well, and that can be improved fairly significantly by putting a "straightjacket" on it (making it use APIs or selecting only tokens that fits a given grammar)
What you are describing sounded decades away last year. Today it's not something you can rely on quite yet if you need accuracy. But if you still believe it's decades away then you need to update your perception of what's going on in AI right now. Forget the free GPT-3.5 or Bard, you really need to try GPT-4 to understand what's about to happen.
Pt sees Dr in office A, has labs and scans ordered. Labs order is printed out and given to pt, who takes it to lab A. Scan #1 is ordered, and order faxed to facility B, who independently calls, schedules, performs, and interprets scan #1, same for scan #2, but at a different facility. Office A has no electronic link to any of these facilities
How does LLM get the data?
A human knows the patterns in the area, and calls around for it, then collates it and feeds in to office A's EHR, where if lucky, LLM can do it's thing
LLM is only saving the bare minimum of the work
In an industry that is still reliant on the fax, I'm not holding my breath for the advent of the LLM savior
What you describe, if it worked reliably, would be an advance over current human-level average performance. And AI will absolutely be able to do it, too. Probably within ten years, though it will take longer to deploy after the capabilities are there.
I just think the y combinator crowd underestimates the human obstacles and ancient tech in the path of getting AI implemented
Also you said the office doesn’t have an electronic link but your example workflow includes the phone ;)
I book a video consult. I talk to a doctor, who has an initial description of my symptoms from my booking on her computer. She takes notes on a computer (I have access to them). If needed, she orders labs on a computer. I go in, and there is no paper, just an electronic record, and while there may well be paper in the process in some instances behind the scenes, data entry people ensures it is all added to the electronic record. If labs are not needed, she'll submit my prescription electronically to my local pharmacy.
This is how it's worked when I talk to my doctor for years (UK; not nearly all GPs have gone this far)
It may well "only save the bare minimum" of the actual labour happening behind the scenes, but with a well functioning system, it saves the expensive, resource constrained part of the work.
* Run database query to see where the type of scan is typically ordered (we know they can generate sql pretty alright)
* * Alternatively someone at one point needs to just make a list
* Call the office (voice conversations is clearly possible, it's a feature in chatgpt as an app)
* Receive fax
* gpt-4-vision or OCR then GPT-4 or however you want to interpret the data
It's not like it's not work to build something like this, but all the pieces are right there.
By stopped cold you mean "there are actual real world self driving cars delivering actual real world passengers in select locations"?
At my GP at least, all of this is presented to doctors digitally, all notes are digital, and all requisition of follow up tests are digital, and my initial consultations are all phone or video and only follow up visits are in person and only if a physical examination or tests are needed.
> Simply very good information synthesizers
> Simple very good syntactic synthesizers
I argue that theorists tend to adhere to the former, and practitioners the latter.
Consider prototyping: For an expert, it costs nothing (or near nothing) to check the output. You even start better, because you know how to ask the right questions to the tool in the first place.
LLMs dont get the semantics right, they get syntactic correlations right. When a subject is common place (literature), the correlations are good enough to get the reasoning right.
There is an IMPRESSIVE amount that gets done with just this. However, expecting to get reasoning right is the bridge too far.
Wrong expectations of tools, result in bad management projects and failures.
You can test this out right now. IF an LLM is able to do information synthesis, it is trivial to set up parallel prompts to work as teams. Try it out.
I did, I know why it can’t work as a result. The simple question to ask is “whats your error rate, and your hallucination rate?”
LLMs are always ‘hallucinating’.
I do hope, that the difference eventually becomes academic. That there is so much training material, that correlation is equal to reasoning. However, it will not be reasoning/semantic prediction.
Finally - the real world work has properties that are emergent. You can read every spec sheet you like, but when you start assembling components together, they are going to do weird things.
What do you mean? I've had GPT-4 (3.5 I think too) talking to itself, setup as a planner and someone critiquing a plan, the end result is much better.
> However, expecting to get reasoning right is the bridge too far.
Hmm, how do you position othello-gpt in this? It builds a world model and makes moves based on that, so it's not just correlation of input (x,y,z typically followed by A so respond A).
See how far that goes, before it’s nonsense talking to nonsense.
2) Production is proof. Get a working LLM based reasoning system that doesnt need its hand held. The process of failing is educational enough.
For the record, I would love it if these things work.
I think it is clear that GPT-4 contains a LOT of information. You can ask it explicit factual questions and it often gets the right answer. When it gets the answer wrong, its wording is typically syntactically correct English, or syntactically correct code.
I'd argue that just because some of the errors it commits are "semantic errors" such as calling a method by an incorrect (but often similar) name in a segment of code or printing false statement in well-crafted prose, it nonetheless gets a lot of the semantics right.
What is reasoning besides a set of language patterns that we define as valid reasoning? Imagine evaluating statements in various formal logics. Nonsense in one can be valid in another based on semantic rules alone.
One could derive the model (model as in model-theoretic semantics) of a formal logical system by sampling a list of valid and invalid statements.
LLMs are doing that kind of thing, it seems. There are gaps, but they are not necessarily gaps that the LLM itself cannot notice.
For example, I will often ask GPT-4 to formulate a plan for something or to create a list of priorities/considerations for an undertaking. I will then ask it to draft an initial plan. After that I will ask it to review/critique its draft based on the initial goals. It typically points out exactly the kinds of gaps that a human would point to as deficiencies that indicate the LLM is not reasoning.
In my view, this indicates that the "knowledge" of how to do the task was always aviailable to the LLM, but the interface (or some aspect of the internal implementation) did not allow the knowledge to be applied all at once. This is not necessarily dissimilar from human intellectual work, in which drafts and self-critiquing is not an unreasonable series of steps.
I have done some work with parallel promopts and various "roles" for different LLM interlocutors toward the same task. While it does sometimes go off the rails, it seems clear that multiple prompts with role-based instructions do achieve a greater level of analytical rigor than a single prompt.
However, it is a TEDIOUS process. It’s simplest to just make something - the parallel prompts scenario for example.
Can you leave your LLMs to their own business, and will they have a workable product at the end of it?
As you said it would seem as if they have the “knowledge” of the task. It should not be an issue.
However you will not get anywhere. The issue isnt in the LLM - the issue is in misunderstanding what is going on, and therefore expectations.
LLMs don’t notice things - in the human manner you assumed they do. It isnt point out gaps, its repeating text patterns.
Our habit of dealing with humans is filling in these gaps and supporting an assumption of ‘noticing’ or ‘improvement’.
I think it is hard for us to consider changes in text, without seeing the changes in meaning as well.
With LLMs you have to accept that it’s not seeing meaning, it’s just seeing correlation.
True. It is important to avoid anthropomorphizing LLMs, etc. I do this intentionally now and then but I agree it is dangerous to do it accidentally.
> I think it is hard for us to consider changes in text, without seeing the changes in meaning as well.
True, but since LLMs were trained on text sequences that had meaning, much of the meaning was accidentally embodied in the resulting model. Areas where LLMs seem to reason well happen to be the areas where the training data was sufficiently generalized and the language tokens used similarly enough... such as recipes.
> With LLMs you have to accept that it’s not seeing meaning, it’s just seeing correlation.
True. I think it is interesting how much it often feels like knowledge.
I think we also have to be careful not to overly glorify human "knowledge" as something other than producing a pattern of output signals in response to a pattern of input signals.
However, it’s taken an inordinate amount of time to achieve even the finesse shared in this comment chain. It’s a challenging topic.
We have to carve out levels of utility for systems that generate output given inputs. Some way to distinguish between what our wetware achieves, and what LLMs achieve - Without putting our models on a pedestal.
(I miss House and hate how it ended)
CDSS: Clinical Decision Support System: https://en.wikipedia.org/wiki/Clinical_decision_support_syst...
Treatment decision support: https://en.wikipedia.org/wiki/Treatment_decision_support :
> Treatment decision support consists of the tools and processes used to enhance medical patients’ healthcare decision-making. The term differs from clinical decision support, in that clinical decision support tools are aimed at medical professionals, while treatment decision support tools empower the people who will receive the treatments
AI in healthcare: https://en.wikipedia.org/wiki/Artificial_intelligence_in_hea...
Med schools tend to try and rejigger their curriculum frequently to try and justifiably make the experience friendlier and w/ less workload for the students, but sometimes, it backfires in terms of this.
I had an older school professor who liked to remark, in his day, no such thing as a differential diagnosis, just the right one and a bunch of wrong ones.
I hear internists talking about +LL and -LL all the time and I'd hope that reflected some higher degree of rationality when reasoning out diagnoses, but perhaps not?