Study: ChatGPT outperforms physicians in quality, empathetic answers to patients
today.ucsd.edu
today.ucsd.edu
> In this cross-sectional study, a public and nonidentifiable database of questions from a public social media forum (Reddit’s r/AskDocs) was used to randomly draw 195 exchanges from October 2022 where a verified physician responded to a public question
They didn’t go to physicians in a patient setting. The physician answers were taken from Reddit threads where they were interacting with people who were not their patients.
Reddit has its own dynamic and people tend to get snarky/jaded. Using this as a baseline for physician responses seems extremely misleading.
-Made to wait 45-60 minutes past their appointment time bombarded with pharmaceutical advertisements
-Spend another 20 minutes sitting in the examination room staring at pharmaceutical company sponsored ads
-Nurse takes your history with a bunch of redundant questions they already have the answers to
-Finally the physician arrives
-Assesses the patient without laying their hands on them
-Ignores everything you say
-Does no tests and cultures no pathogens
-"take some antibiotics and I'll bill your insurance company $500, see ya"Then when the patient asks ChatGPT if the tests could give false negatives, it could provide some very valid-sounding answers that say repeat testing might be necessary.
This isn’t a hypothetical. It’s currently happening to an old friend of mine. They won’t let it go because ChatGPT continues to give them the same answers when they ask it the same (leading) questions. At this point he can get ChatGPT to give him any medical answer he wants to hear by rephrasing his questions until ChatGPT tells him what he wants to hear. He’s learned to ask questions like “If someone has symptoms _____ and ____ could they have <rare disease> and if so how would it be treated?” A real doctor would see what’s happening and address the leading questions. ChatGPT just takes it at face value.
The difference between ChatGPT and real doctors is that he can iterate on his answer-shopping a hundred times in one sitting, whereas a doctor is going to see what’s happening and stop the patient.
ChatGPT is an automated confirmation bias machine for hypochondriacs.
The subtext you're missing here is that GPT with access to the entire corpus of medical data could undermine the entire money printing machine (referring to US healthcare here). What test would cost thousands of dollars if the only human cost to run it is drawing some blood and putting it in a machine?
Happened to my granma when she was 90. She had convinced her doctor prescribing her 5 or 6 drugs. She swallowed them until she was so weak she had to go to hospital, where the tests revealed that she was 100% sane. She didn't like it.
> they do the required tests and examinations
Do you have a source or is this anecdata? I do agree with your other claims though.
-“take some antibiotics and our billing dept will bill your insurance company for the maximum amount they estimate your insurance can pay, plus some margin, see ya”If you’re a nurse and your job is to take histories, it’s better to take the history in the same way every time, systematically. This minimizes the chance of making mistakes.
Moreover the notes may be wrong, or you may give a different answer this time around. You might give an important detail this time around that you didn’t give the last several times, which actually turns out to be consequential.
It might seem like a waste of your time but it’s really not. Measure twice cut once.
Citation needed. For example, patient may get bored and only answer the first few questions accurately each time.
Canada's various provincial public health care systems (with a corresponding lack of private health care offerings) tend to be like that.
Many Canadians don't have a dedicated physician. Even if they want one (or want a new one), it's often difficult, if not impossible, to find one who's close by and who's accepting new patients.
Having a dedicated physician still often results in an experience much like that other commenter described. It's not an exaggeration. Long waits even with appointments, rushed examinations, and low-quality service are the norm.
Another option, which is sometimes used even by people who have dedicated physicians, is a walk-in clinic. Unfortunately, they can be quite rare and inconvenient to get to, even in Canada's largest cities, assuming they're even open when you need them. You'll usually face an even longer wait, even less time with the doctor, and typically see a different doctor if any sort of followup is needed.
Then there are hospital emergency rooms. That usually means getting to the nearest sizable city, and even once you're there, you've got to be prepared to wait many hours, even for relatively serious situations.
Ultimately, in Canada, it doesn't matter whether you see your family doctor (if you even have one), use a walk-in clinic, or go to a hospital emergency room. It's going to be a horrible experience, and there's pretty much nothing the average person can do about it.
Given the lack of competition and due to other government-imposed market distortions, there's no incentive for doctors to offer anything resembling good service to the general public.
The best situation is to have a doctor who's a close friend or family member, and who may be able to help mitigate at least some of the typical problems.
The next best option for Canadians, assuming they have the money for it, is often to seek treatment in the US or overseas.
Just "finding a new doctor" isn't feasible, unfortunately.
- Your 40 minute in-depth appointment is finished in an amazing 5 minutes
- They can do this amazing appointment time compression because of an encoding technique called "one size fits all"
- The physician's assistant is so efficient that they've already scheduled your follow-up appointment - for three months from now, because they know you just love the suspense.
The potential seems there: cash payments accepted, many convenient locations, possibly lower wait times.
Of course CVS is Aetna, for better or for worse.
To be fair, they were so heavily booked that they weren't taking new patients and you had to wait weeks or months for non-urgent visits
Also some of them went out of business..
Meanwhile my mom has had cronic pain without any proper diagnosis. However, the German public health care system must have spent €50-100k in tests and to my astonishment she is currently in a clinic for 3 weeks that focuses on undiagnosed pain. As close to Dr House as it gets.
That’s not a killer argument in a comparison with an AI.
Palpation, especially for musculoskeletal issues, is an incredible diagnostic tool. And manual therapies are surprisingly often an effective alternative to surgery or drugs.
A friend of mine, who is a primary care physician, just quit her job after the clinic changed its policy so that she had to fill her schedule with 10 minute appointments (6 per hour) and double book at least 10% (because the clinic was losing SO much money to no-shows /s). She burnt out after 6 months of this and moved somewhere they give her a more reasonable 15 minute time slots and no double-booking…
All this is to say, it may be a doctor thing, but I would bet it’s more likely a clinic trying to squeeze as much money as possible, and the doctor getting none of the extra profit from it.
BTW this is done to verify the records and make sure you are who they think you are. Giving someone another persons treatment can be… bad.
Is this a regular occurrence or have you found a better physician since then?
You forgot at least mildly insulting the patient by calling him stupid.
- remove your clothes and wait in a freezing room for 30+ minutes
(Largely a compliance exercise to reinforce the social status hierarchy.)My experience with most doctors is that they are among the least empathic people I have ever dealt with. I think that using AskDocs actually gives doctors an unrealistic _advantage_ in the study.
It's apparent, at least in the US, that a lot of GPs can be unhelpful but believing that accurate diagnosis via text without providing evidence, or going through some type, of physical examination is feasible by Reddit expert docs, only demonstrates lack of critical thought.
Do you know if the subreddit does anything to address that?
It’s almost always the headlines and PR pieces that exaggerate it.
“ChatGPT is more empathetic than Reddit doctors” isn’t interesting. Strip the “Reddit” out and then everyone can substitute their own displeasures with doctors and now it’s assumed true.
That is definitely happening here.
https://jamanetwork.com/journals/jamainternalmedicine/fullar...
> evaluators did not assess the chatbot responses for accuracy or fabricated information
Yikes. (I do fault the researchers quite a bit for quietly slipping that little detail into a page-long "limitations" section.)
So we have here "ChatGPT performs better than doctors *"but* we sampled on an unpaid internet forum"
The other one thats doing the rounds where noone is verifying the source is "ChatGPT develops lethal chemical compounds *but*" It would be too unethical to verify with an actual chemist if any of these compounds it dreamt up actually do anything".
ChatGPT is showing how bad the economics of journalism are for understanding a topic outside of taking press releases at face value*
For my experience is terrible for any question I ask: - if it doesn't know well a person it simply makes up things - if I ask code examples or scripts, most of the time they are wrong, I need to fix them, they contain obsolete syntax etc... - if I'm asking a question, I'm expecting being asked for more context if the subject is not clear, instead it starts spitting text without even realising I asked a completely different thing etc...
I could go on for hours with other examples, but I'm seriously not finding it useful
Some times this presenting of problem to it means I spend anywhere from 5-10 mins actually writing the points down that describes the requirement - which would result in a working component/module (UI/backend).
We have been trialing GPT4 in my company and unfortunately almost everyone's experience is more on the lines of yours than mine. I know it shouldn't, but honestly it frustrates me a lot when I see people complain that it doesn't work :). It definitely works but it depends on the problem domain and inputs. Often people forget that it has no other context about the problem than just the input you are providing. It pays to be descriptive.
It's almost like "AI hacking people's brains" turned out to happen accidentally, and a huge number of supposedly smart people are getting turned into mindless enthusiasts by nothing more than computer generated bullshit.
I only believe if they actually trained chatgpt on those type of tests specifically.
Not the actualy dynamic nature of dealing with patients & lawsuits.
NLP can recognize alt accounts of individuals on places like HN and reddit, but a person would probably need to study the comments pretty hard to determine the same thing, its not natural for people imo but it seems to be the foremost aspect of any kind of model that's processing human writing.
I'll also note that when there was hype around Stable Diffusion, one of the images shared around was that of an astronaut riding a horse. If you actually run Stable Diffusion with its default tuning and ask for that prompt, you will get 6 images, of which 5 of them are outright disasters (horses with 6 legs and going downhill from there), and then the 6th image, the only one which could possibly pass as a decent result, is the one that everyone shared and reshared and hyped. Usually other prompts give even more terrible results where there are 0 passable images without extensive tuning. Stable Diffusion now is acknowledged to actually be crap --despite the hype-- and I supposedly need to try the next best thing, whatever that is. But I find myself facing the same situation with ChatGPT 3.5, and now with ChatGPT 4, despite the fact there is no "next best thing", and I don't even know how they could even possible try to fix the problem of it being just wrong.
I do agree that it's bad at following the prompt exactly. It will produce most of the things you mentioned in the prompt but not necessarily in the same fashion you asked for. I don't agree that produced images are mostly shitty, just visit that subreddit.
SD is definitely very good in the right hands, and it’s a little unfair to expect to be able to get instant good results without any skill. It’s honestly pretty crazy that we now have things like ChatGPT and SD – and people are already calling them crap because they don’t work perfectly and their productive use actually requires some skill!
But r/StableDiffusion, or any public gallery, is obviously one giant selection effect. 99.9% of attempts could be crap and the 0.1% would still be enough to fill a subreddit.
To give you a general idea of what percentage can be good images: I recently made a lots of wallpapers to cycle through daily using SD. I found a good prompt, a good model, and let it generate bunch of images continuously for few hours.
None of the images were shitty (they were all random seeds), only images I discarded had artifacts I didn't like or couldn't keep my eyes away from. With SD you can't just expect to give a prompt "beautiful landscape" and expect it to give you a beautiful landscape. It won't. You shall get shitty images and might get a few pleasing ones. You must tune your prompt to get good results.
I use it for my side projects, for tech I have no experience in, and it works very well, because I know what I want, I know that it is possible and I just need it to vomit the boilerplate to save me 5 google searches
For my day job it's next to useless, and if your day job can already be automated by chatgpt I have bad news for you
If they can automate the work of a physician, who exactly is safe? Low skill labor, maybe, for awhile.
(This is a genuine question) Is it hard to find a solution to that question on the web by using a search engine?
Advanced users of both Midjourney and SD can get some stellar results out of them. Some of that is due to trial and error, and going through dozens or hundreds of images to pick the best ones, but being adept at crafting prompts and using other features of the programs plays a big role too.
Know your tools.
I think the coined phrase is "prompt engineering".
Side note, where's the eye roll emoji hiding?
goalposts
To present the best prompt/response while not disclosing that it is a result of trial and error is a different thing altogether.
Fixed that for you.
If you have very little technical aptitude, patience, and willingness to learn, it would be better for you to stick with some thing a little more newbie friendly such as mid journey.
However, loading a proper model such as 526, realistic Vision, or ReV, using control net, and guiding a prompt to where you want it to go while making changes using inpainting/img2img/etc can result in stunning images.
Obviously you have to play around with them and figure out how to prompt & tune them.
Once you do that, you can get pretty amazing results.
I make do a bunch of social media graphics and stable diffusion + midjourney has been insanely useful.
You literally cannot respond with "you are holding it wrong" specially when I'm claiming that even for the popular _example prompts_ SD authors used they had to hand-pick the best random result over a sea of extremely shitty images.
And even in the original paper they disclaim it by saying "oh, our model is just bad at limbs". No, it's not just bad at limbs. They just happened to try examples where it could particularly show how terrible it is at limbs (i.e. spider legged horses and the like). But in truth, it's just bad at everything.
Besides, there are models that are much more capable than Stable Diffusion. The best one currently seems to be Midjourney V5.
I'll repeat myself. You have to play around with the models and learn how to use it (just like you have to do for everything) .
> But in truth, it's just bad at everything
Thousands of people (including myself) have had the complete opposite result and have gotten amazing pictures. You can play around with the finetuning with different models from civitai and get completely different art styles too.
Like, this is so dumb I don't even know how to respond lol.
You're like some guy who got a computer for the first time and couldn't figure out how to open the web browser, so he just dismissed it as useless.
I do a lot of my SD with a fixed seed and 1-image batches, once you know the specific model you are using getting decent pictures isn't hard, and zeroing in on a specific vision is easier with a fixed seed. Once I am happy with it, I might do multiple images without a fixed seed using the final prompt to see if I get something better
If you are using a web interface that only use the base SD models and don’t allow negative prompts, yes, its harder (negative prompts and in particular good, model specific, negative embeddings are an SD superpower.)
I do agree that it's bad at following the prompt exactly. It will produce most of the things you mentioned in the prompt but not necessarily in the same fashion you asked for. I don't agree that produced images are mostly shitty, just visit that subreddit.
This doesn't really say anything, because it's just survivor bias. The entire purpose of that site is to show the successes. Most people get disasters every single day, they just don't upload them to the site. Even if I try the same prompt I will get a disaster image, as long as I don't use e.g., exactly the same random seed they happened to use. This is not even "prompt engineering". It's just outright playing with the dice.
> You need to agree to share your contact informations to access this model
All of this does feel like another scam where the tech is exaggerated and side hustles are actually the endgame.
If, on the other hand you actually put together a prompt which tells it what you expect, the results are very different.
E.g. I've experimented with "co-writing" specs for small projects with it, and I'll start with a prompt of the type "As a software architect you will read the following spec. If anything is unclear you will ask for clarification. You will also offer suggestions for how to improve. If I've left "TODO" notes in the text you will suggest what to put there." and a lot more steps, but the key element is to 1) tell it what role it should assume - you wouldn't hire someone without telling them what their job is, 2) tell it what you expect in return, and what format you want it in if applicable, 3) if you want it to ask for clarifications, either ask for it and/or tell it to follow a back and forth conversational model instead of dumping a large / full answer on you.
The precise type of prompt you should use will depend greatly on the type of conversation you want to be able to have.
Like I know ChatGPT-4 can generate a bunch of code really quickly, but is coding in python using pretty well known libraries so hard that it wouldn't just be easy to write some code? It's super neat that it can do what it does, but on the other hand, modern editors with language servers are super efficient too.
My point was that if you just ask it an ambiguous question it will return something that is a best guess. It's what it does. To get it to act the way the person above want it to, you need to feed it a suitable prompt first.
You don't need to write a new set of instructions every time. When I "co-write" specs with it, I cut and paste a prompt that I'm gradually refining as I see what works, and I get answers that fit the context I want. When I want it to spit out Systemd unit files, I cut and paste a prompt that works for that.
The stuff I'm using it for is stuff where I couldn't possibly produce what it spits out as productively not because it's hard but because even typing well above average typing speed I couldn't possibly type that fast.
An intelligent agent shouldn't need this type of prompting, in my opinion.
There are many categories of usage for them, and relatively "dumb" completion and boilerplate is still hugely helpful. In fact, probably 3/4 of my use of ChatGPT are uses where I have a pretty good idea what it'll output for a given input and that is why I'm using it, because it saves me writing and adjusting boilerplate that it can produce faster. Most of the time I don't want it to be smart, I want it to reliably do almost the same as it it's done for me before, but adjusted to context in a predictable way (the reason I'll reach for it over e.g. copying something and adapting it manually).
We use far dumber agents all the time and still derive benefits from it. Sure it'd be nice if it gets smarter, but it's already saving me a tremendous amount of time.
It’s a save 15 minutes here, 20 minutes there kind of thing that can add up to hours saved over the course of a day.
> if it doesn't know well a person it simply makes up things
Asking it for factual information about a subject can be a bit hit/miss depending on the subject. Better to use bing chat, because it will use info from the web to inform the response
> if I ask code examples or scripts, most of the time they are wrong, I need to fix them, they contain obsolete syntax etc...
How wrong? More wrong than having a junior or mid level developer contributing code?
Think about it a different way: you just gained an assistant developer that writes mostly correct code in seconds. Big time saver.
Also: if you want it to use a particular code style etc, give it few shot examples.
> if I'm asking a question, I'm expecting being asked for more context if the subject is not clear
Then you need to tell it that in you prompt: "if the subject isn't clear, ask me some clarifying questions. Don't respond with your answer until I have answered your clarifying questions first". Or: "ask me 3 clarifying questions before answering" to force it to "consider" how well it "knows" the subject first.
ChatGPT isn't an AI in the sci fi sense of the word. It's a language model that needs to be prompted the right way to get the results you want. You will get a feel for that the more you use it.
This is disingenuous. OP is right. ChatGPT is mostly inaccurate and contextless by nature.
It always produces something confidently so it's very easy to think it's the right answer.
Now, it can be extremely useful for any task that can be easily verifiable and you don't know the syntax, how to approach something, etc. Because any decent software developer can use it as part of the prototyping or brainstorming and get wherever they want to go in coordination with ChatGPT.
What you can't do is assume it's "the expert". You are the expert and you're the intelligent one and the chat generates potential useful things.
That's on verifiable stuff. On other things it can be so laughably bad it's impressive how much "everyone" (as in articles, hype, HN users who downvote criticism as being from luddites) pushes it as something that is not: an AI that can think and be relied upon and any inacuracy will just "get better with time" aa opposed to "the model and also the politics around it" doesn't have a path to get to that idealisation that is being sold.
It's a tool, but it's usually sold as better than it is, like in this case, presumably with the intent of relying upon it to save cost in some key integration point. The problem is that it mostly won't work and comparing it to bad humans or worse integrations doesn't show the fundamental low ceiling that it has.
I think people with authoritarian mindsets (and I don't mind Left or right but inherent trust in authority) easily want this to be a source of truth they can magically use, but there's no path for that to be true. Just to appear true.
Yes. The failure mode of a human and chatgpt are nothing alike — I am far more experienced in spotting beginner mistakes in code reviews, then seemingly good, but actually illogical bullshit code generated by LLMs.
I have never had it produce non-trivial, novel code that was correct, so I mostly use it as a search engine instead.
Either that or the problems they're solving are nothing boilerplate couldn't handle.
The article compares verified responses on r/AskDocs (yeah, a subreddit) and those from ChatGPT. That's it. How is its coding compatibility even remotely relevant? It's like saying "Excel is bad in editing photos, so it must be a bad spreadsheet software as well."
If you ask it what the facts are, it just gives you a load of nonsense.
That said, I can’t really use ChatGPT as a search engine, but I did plug it into a self-hosted telegram bot and I do ask it some basic questions from time to time - telegram is a good UI for it.
I usually need someone to explain something to me. And I used Google before to land on a site where I could find the explanation (e.g. how to use a library). ChatGPT can explain most things I need and I can skip Google and the other sites. But it's indeed not a search engine, if you need factual information then your best bet is to find the documentations, articles, databases, etc.
Put another way: People who don't pick up these tools and learn how to be effective at them will increasingly be at a significant performance disadvantage from people at their skill level who do pick them up.
How do you know if they don't? Do you expect them to report back to a throwaway comment on HN? The sentiment is shifting on HN, that much is clear, just like it shifted with crypto.
It seems like you could maybe automate that. Let it spit out its first draft of an answer, have the framework tell it "please correct the errors" and then let it have another go and only present the second attempt to the user.
I'm reminded particularly of a screenshot of someone gaslighting ChatGPT into repeatedly apologising and providing different suggestions for Neo's favourite pizza topping, despite it answering correctly that the Matrix did not specify his favourite pizza topping first time round, but it applies equally to non-ridiculous questions
I find it very helpful to ask a series of questions and see a number of examples to get a primer on what to expect with something. The main benefit over Google or going straight to the docs is I can start with my specific requirements. I then dig into the documentation to deepen my understanding. I can typically move forward with ChatGPT generating some code as a starting point.
It can be incorrect or out of date but combined with my experience I find myself being more productive with it.
A weakness I see is complex code requirements. It knows what it knows.
I note that you seem a little frustrated with vague or incorrect responses. It helps to tell ChatGPT the role it should play. It helps as well to instruct it to ask questions of you to improve the response. Personally I prefer to tell it keep its answers brief, I get less walls of text and I can narrow in on the specific answer I am after more quickly.
After reading the article, what you wrote doesn't seem to make much sense.
N.B. as well: If someone thinks they have bleach in their _eye_ and can still open their eyes enough to write a Reddit post, much less read through ChatGPT’s extremely long answer, they’re almost certainly fine.
It costs you zero dollar to not post an irrelavant comment then. I really wonder how you justify this "I didn't read the article, but I have a very, very strong opinion (straight up calling it a made-up) on it" behavior. The internet is rotting people's brains I guess.
Yes. Down in the limitations section of the study (https://jamanetwork.com/journals/jamainternalmedicine/fullar...):
> evaluators did not assess the chatbot responses for accuracy or fabricated information
That is a... significant issue with the methodology of the study.
- write a contract for the sale of my motorcycle: put all details, names and numbers with labels on a spreadsheet, paste on the chat and ask for a contract, then edit.
- learn french: I told gpt "when I write wrong stuff in french, always let me know and teach me the correct ways". Then, after a few weeks I asked for a .csv with the stuff that he corrected so I could import into Anki, which actually worked.
- coding on a daily basis: I am learning Rust on my new job, so I ask it things all the time, it helps me a lot.
I both use ChatGPT to boost productivity but also see the amount of mistakes it makes and will keep making and am surprised at the extreme denial of anyone who tries to shut down criticism of the wrong type of hype (the one that sells something that is not there)
Thanks for such a clear cut example for my points.
Would you be willing to publish the contract with sensitive information redacted?
All these bold claims are just claims until people come up with some substance. Talk is cheap and confirmation bias happens all the time.
The point still remains: let's see what the LLM delivered that the user actually used. Either it's legally binding and an appropriate use, or it's not fit-for-purpose.
Equally, why not an interactive form using conditional logic? No hallucination possible. Much more simple and reliable.
If you know this, you don't need GPT.
If you don't know this, you don't have a way to assess GPT's attempts at a contract. A bill of sale is indeed simple, but there's a lot of more subtle legal issues someone might run into in life.
If you ask it for boilerplate or for something that's a basic combination of things its seen before, it can give you something decent, possibly even useable as-is. But as soon as you step into more novel territory, forget it.
There was one case where I wanted it to add an async method to an interface as a way of seeing if it "understood" the limitations of covariant type parameters in C# with regards to Task<T>. It did not. I replied explaining the issue and it actually did come back with a solution, but it wasn't a good solution. I told it very specifically that I wanted it to instead create a second interface for holding the async method. It did that but made the original mistake despite my message about covariance still being within the context fed back in for generating this response. I corrected it again, but the output from that ended up being so stupid I stopped trying.
And at no point was it actually doing something that's very important when given tasks that are not precisely specified: ask me questions back. This seems equally likely to be a problem for one of these language models replacing a doctor. It doesn't request more context to better answer questions so the only way to know it needs more is if you already know enough to be able to recognize that the output doesn't make sense. It basically ends up working like a search engine that can't actually give you sources.
E.g., lots of HN users claim to use it for dev or learning new programming languages. Given the frequency of hallucination and their Dunning-Kruger complexes in full-effect, they don't know when it's teaching bad information or functions that don't exist.
It's an LLM. Not an AGI.
One reputed physician completely refused to acknowledge that a specific medication might be causing some of my side effects, even when I shared links to peer reviewed studies from reputed universities (including his alma mater) that specifically talk about side effects from the medication.
My success rate with doctors is about 20% at this point. I've stopped visiting them for most ailments, and if I ever do get any prescriptions, I make sure to research them thoroughly. The number of doctors who will casually prescribe heavy duty drugs for common ailments is unreal.
Combined with the prestige of the profession, many turn into egoistic know-it-alls whose real competence is equivalent to a car mechanic who tells you you're just driving your car wrong when you come in with anything that's not obvious from a 15s inspection because they get paid anyway.
I would be surprised if most of the doctors I've seen could outperform even GPT-2.
Even so doctors are extremely overworked now and insurance companies don't want to pay their fees. So now they're running from patient to patient, unable to even pay attention to them. That's even if you get a doctor now before the multiple nurse practitioners or physician's assistants who also don't know anything.
I’ve had a couple of surgeries and the eagerness of doctors to give me opiates for pain relief was baffling, even when I clearly told them that I can tolerate the pain and don’t need anything stronger than ibuprofen.
Also opiates in a short term setting are good meds. Pain control is good and people are able to get moving faster.
There are good doctors out there but there are a lot of bad. I always advocate people if they're unsure to get a second opinion. Just like you would if a plumber said "you need to replace the whole system." If it doesn't seem right or you don't feel like you got the proper attention, go somewhere else and see.
Medicine isn't an exact science for much of us. It's a lucky thing when it's a simple infection that antibiotics can cure. Most of our problems aren't so easy. Just be slightly skeptical. Don't go "Fruit will cure my pancreatic cancer" crazy either.
Buried deep in the limitations section:
> evaluators did not assess the chatbot responses for accuracy or fabricated information
Repeat:
> EVALUATORS DID NOT ASSESS THE CHATBOT RESPONSES FOR ACCURACY OR FABRICATED INFORMATION
https://xkcd.com/937/ comes to mind. It's not implausible at all to me that ChatGPT could outperform Reddit in detail/manners for health advice (and honestly, even for actual doctors I've heard some horror stories about bedside manner and refusing to actually believe/consider symptoms), but if the study isn't actually checking that, if they're just checking if the chatbot was more polite/empathetic... that's a huge qualification that should be up-front and center.
His response to my (limited and naive) question was essentially, people will hold attachments to beliefs about the human element in these transactions, and the philosophical part of the question was why people would hold beliefs that were not strictly rational, and he sort of declined the implied ethical questions. There was no reason to expect him to respond with more, but given I was navigating the ethics of AI driven medical care via questions of privacy in system design decisions (read: accountability) for the institutions who would use it and the millions of people subject to it, it seemed like an opportunity to be at the very forefront of what would likely become a defining social issue of our lives in a couple of short decades.
What we discovered then as architects, as most people are just about to, is that the main use case for ML/AI will be to use complexity to diffuse accountability away from individuals, and aggregate it up into committees and ultimately corporate entities that are themselves artificial beings. AI is the corporeal manifestation of an institution, essentially a golem.
ChatGPT has been helpful for fairly in-depth research, as well as answering questions I don’t want to bother my oncologist with.
Actually, I prefer it to be less empathetic and more robotic with its answers. When it tries to be empathetic it goes beyond what I’ve asked it and sprinkles on typical safety filter bs.
In contrast my oncologist could be a bit more empathetic. This person directs your care, and if they do so in a way that seems like a black box it can be rather scary. They know enough about me, with all of their experience, to be a “helpful human” in a difficult situation.
I’ve have yet to experience any of the domain knowledge I have being incorrect.
Similarly, today I saw this [0] comment and in an attempt to help the user went to ChatGPT.
Prompt:
> User I remember in the 2000s, there was some site using a decimal IP address (as a single number, not dotted quad) that had hacking/crypto puzzles. Something with a "Alice in Wonderland" theme. Does that ring a bell for you?
Response:
> I'm sorry, but as an AI language model, I don't have personal experiences or memories. However, based on your description, it's possible that you might be referring to the website known as "Project Wonderland." It was a website that featured a series of cryptographic puzzles based on the story of Alice in Wonderland, and it used a decimal IP address (203.208.46.147) as its URL. The website was active in the early 2000s and gained a cult following among enthusiasts of cryptography and puzzle-solving. However, the website is no longer active today.
I got really excited to have found an answer until through Google and the Wayback Machine I realised that ChatGPT just made this all up and no such website existed at that URL.
I tried starting a new chat with ChatGPT to ask it about this "Project Wonderland" website and it had no idea what I was talking about.
[0]: https://news.ycombinator.com/item?id=35748714
(I am using ChatGPT 3.5 (March 23, 2023))
The server might even have refused connection if the HTTP "Host: " header wasn't in that format, but as a human, rather than a language model, I'm not sure about that and might be misremembering :)
* 3301 is one of three numbers that had to be added to get the .com url * The Wikipedia page cites someone close to the group with "follow the white rabbit" * Years don't quite match up but given that you only asked if it rang a bell, that is fair enough
But then I said No, I'm talking about JR in babylon 5
The response was
I apologize for the confusion. "JR" in Babylon 5 refers to the character John J. Sheridan, who was the captain of the titular space station during the later seasons of the show. Sheridan was never refereed as JR in Babylon 5. There was no JR in Babylon 5. But because I asked, ChatGPT happily made up things.What I got back was confidently incorrect, and gave me examples that were neither of the two. (Interestingly, most of the examples were chiasmus instead.) I tried correcting it to get better responses, but it kept apologizing and then giving me the incorrect examples back.
Ten to fifteen years of doctors, over half a dozen, failed to accurately diagnose her until this month.
ChatGPT got her diagnosis right in a second. Wow. I'm both amazed and angry that we could not get this diagnosed a decade ago.
It could also be that knowing the correct diagnosis changes your description to highlight things in a way that suggests the correct outcome. I believe doctors are also susceptible to that effect.
I'm not trying to imply that doctors should not do a better job diagnosing. People should not have to "learn" to find the right doctor, or how to operate them. It's a crying shame that people like you have to go on years-long journeys to get correct help, and I'm sorry that you and your wife had to go through that.
Just saying it might be slightly more apples-to-apples to compare ChatGPT's performance to the last couple of physicians you saw, and not the whole lot of them. But again, that's still a very favorable comparison for the non-human.
Today we do not have feedback to the automated systems. If we start doing following on a large scale Measure->treatment->measure again->adapt treatment, then the system will learn and will make connections that no doctor has ever made, because the minds of all doctors are not interconnected, at least not in a structural manner.
You can chat messages on WhatsApp, SMS, and now a ChatGPT plugin to query a vector database of health information (Pinecone) and respond with GPT-4 for higher quality results than default ChatGPT. The chatbot then prompts the person messaging to verify results with a real doctor and offers to connect them to a Doctor that I'm working with.
I was doing fieldwork and would get sick in foreign countries where I didn't know the language. My friends would drag my food-poisoned body to a hospital where they inevitably would try to give me prescriptions that I wasn't familiar with and wanted to double check. I wanted to build for myself a WhatsApp bot that I could text while I travel to verify health information in low-bandwidth internet situations.
I shared it with some friends, and they shared the WhatsApp contact with other friends and family, and now it's being used by people in about 10 countries around the world, in several languages.
Would love any feedback if you try it out! The phone number for SMS/WhatsApp is +1 402.751.9396 Or link to WhatsApp if easier: https://wa.link/levyx9
But also it can be dangerous when you don't have access to medical information. The first friend that started testing the WhatsApp bot lives in the Sinai desert in Egypt where it's really hard to get to a clinic to ask questions. It's kind of similar in rural Nebraska where I grew up. We're taking things one step at a time and trying to provide the best services that we're able.
It was mind blowing: it identified 2 possible explanation that were already on my radar, 3 more that I had never considered of which one seems very likely and I am currently getting tested for, explained how each of those correlated with my symptoms and medical history, and asked why I had not had a specific marker tested (HLA b27) that is commonly checked for this type of disease (and indeed, my doctor was equally stumped - he just thought that test had been done already and didn't double-check).
Bonus: I asked if the specific marker could be inferred from whole genome sequencing data (had my genome sequenced last year). He told me which tool I could use, helped me align my sequencing data to the correct reference genome expected by that tool, gave me step by step instructions on how to prepare the data and tool, and I'm now waiting for results of the last step (NGS data analysis is veery slow).
Medicine, and in particular diagnosis, is particularly difficult due to the width of knowledge required due to the span of possible diseases.
It completely makes sense that GPT, or similar, would simply be better than doctors at diagnosis in time, and it is very plausible that the time is now.
This is fantastic news for humanity. Humanity gets better diagnosis and we don't put high IQ people thought a grinder that is medical school and residency to do a job which is not suited to human cognition.
It's an amazing win.
""" Here is the patient history (in french):
• 2000-2010:
◦ ...
• 2013: ...
...
Additional Notes:- <Some additional observations and comments - patterns I noticed, family history, etc.>
Patient is particularly worried about xxx and yyy. What are some possible causes explaining these symptoms and the overall history? Give detailed reasoning to support your hypotheses, include differential diagnosis, think about rare diseases (common causes have already been considered), consider possible combinations of diseases, and don't hesitate to ask follow-up questions to improve diagnosis. Please answer in English.
Additional test results that are outside of normal ranges:
- <list of abnormal results>
Try to consider and explain as many of the blood tests in your work up, and ask for any missing information if necessary. """
The key was to address it as if I was a doctor asking for an opinion. If I asked it "as a patient", I got much lower quality and dumbed down answers.
Pride goeth before the fall. I wonder how many arrogant jerks will be humbled to see that they too are now inferior to a computer ("soon"). Humans will always be better at being human though, perhaps they will learn empathy is more important than they thought.
What do you mean by that? If you mean humans will be always the best being the creature we call human then that goes by definition.
If you mean humans will be always more compassionate/emotionaly understanding/better suited to deliver bad news then I am afraid that is unsubstantiated.
ChatGPT = artificial sociopathy
In this case, the existing documentation of such things combined with the events in the GP's own medical history have been fed into a machine that can identify patterns that a human doctor should have, but for whatever reason has not, identified.
I think the potential ramifications for this are huge.
"... about 7 percent of Americans are positive for HLA-B27 but only 5 to 10 percent of people with a positive HLA-B27 will have AS" https://creakyjoints.org/about-arthritis/axial-spondyloarthr...
If I ask GPT4 about some arcane math concept it’ll wax lyrical about how it has connections to 20 other areas of math. But it fails at simple arithmetic.
One of my better math professors in a very good pure math undergraduate program added 7 + 9 and got 15 during a lecture, that really doesn't say anything about his ability as a mathematician though.
Who knows, OP could be a paint sniffer and that’s their root issue. Brainstorming these things requires creativity and even hallucination. But that’s not what doctors do.
Does it though? When allowing LLMs to use their outputs as a form of state they can very much succeed up to 14 digits with > 99.9% accuracy, and it goes up to 18 without deteriorating significantly [1].
That really isn't a good argument because you are asking it to do one-shot something that 99.999% of humans can't.
The cited paper covers this to some extend. Instead of asking the LLMs to do multiplication of large integers directly, they ask the LLM to break the task into 3-digit numbers, do the multiplications, add the carries, and then sum everything up. It does quite well.
If allow LLMs to do the same instead of producing the output in a single textual response, then they will do just fine according to the cited paper.
Average humans can do multiplication in 1 step for small numbers because they have memorized the tables. So can LLMs. Humans need multiple steps for addition, and so do LLMs.
The only reason failing at basic arithmetic indicates something when discussing a human is because you can reasonably expect any human to be first taught arithmetic in school. Otherwise, those things are hardly related. Now, LLMs don't go to school.
If that's your best argument, you don't have an argument.
Literally the majority of the page is basic arithmetic, mostly Bayes. Diagnosis is a process of determining (sometimes quantitative, sometimes qualitative) the relative incidences of different diseases and all the possible ways they can present. Could this be X rare virus, or is it Y common virus presenting atypically?
That is literally true of every marker in existence. It's not the most specific marker, no, but if you already have a strong prior for presence of autoimmune disease, then the presence or absence of that HLA subtype can point towards most likely root cause (autoimmune diseases are all incredibly similar in the early stages).
Contact me at Drgpt@altmails.com
The thing feels human, until you ask it for more options 10 times in a row. Or ask it for a more concise version 5 times. Then it shows: It just doesn’t get tired of your S.
> [...] a roadmap for how to get a medical large language model-based system regulatory cleared to produce a differential diagnosis. It won’t be easy or for the faint-hearted, and it will take millions in capital and several years to get it built, tested and validated appropriately, but it is certainly not outside the realms of future possibility.
> [...]
> There is one big BUT in all this that we feel compelled to mention. Given the lengthy time to build, test, validate and gain regulatory approval, it is entirely possible that LLM technology will have moved on significantly by then, if the current pace of innovation is anything to go by, and this ultimately begs the question - is it even worth it if we are at risk of developing a redundant technology? Indeed, is providing a differential diagnosis to a clinician who will already have a good idea (and has available to them multiple other free resources) even a good business case?
That aside, the training isn't blind, it's guided, and it's likely they use verified correct sources of info to train for some things, like medical diagnoses.
You may also be interested in Apendix A in the same document: "Details of Common Crawl Filtering"
The quality of this study is so incredibly poor that I'm flabbergasted at UCSD's bar for platforming such garbage.
A concern with something like this though, is to what extent is ChatGPT just telling patients what they want to hear, as opposed to what they need to hear?
I recall in a chat a person getting rather exasperated with a coworker and the person using ChatGPT to generate a friendly/business professional "I don't have time for this right now."
GPT could be used in a similar manner - "Here is the information that needs to be sent to the patient. Generate an email describing the following course of treatment: 1. ... Stress that the prescription needs to be taken twice a day."
The response will likely be more personal than the clinical (we even use that word as an adjective) response that a doctor is likely to give.
> ...the original full text of the question was put into a fresh chatbot session, in which the session was free of prior questions asked that could bias the results (version GPT-3.5, OpenAI), and the chatbot response was saved.
It seems like they just pasted the question in. For those who have asked it for medical advice, how did you frame your questions? Is there a prompt that will help ChatGPT get into a mode where it knows it is to provide medical advice? As an example, should it be prompted to ask follow up questions if it is uncertain?
Any kind of LLM and its frequency of hallucination means that LLMs are inappropriate as a solution in this scenario. An LLM is not a physician, it's an LLM.
You can make soup in an electric kettle, but it doesn't make it the right tool for the job, and comes with a lot of compromises.
Somehow I think that the only profession in the US that still uses fax machines will be slow to take up this new technology.
That guy would have failed against ChatGPT, and I loved the way he told things. Anythong else would have just driven me crazy, maybe to the point of looking for a different doctor.
So I giess, what passes as good bed side manners for doctors largely depends on the patient. By the way, the dentist I have since is in the same category as my, luckily former, oncologist. A visit with hom usually takes no more than 5 minutes if he's chatty, less if not. Up to 10 when treatment is required, anuthing longer than thaf is a different appointment.
Not Bing Chat. She doesn't have problems telling users that they have been bad users.
Fun times.
ChatGPT did not actually answer patients.
" Doctors are not trained to respond empathetically to social media posts in public forums. As a cardiologist if I had to respond to such posts online I’d be more worried about liability and saying something w/o full understanding of the patient’s situation. Empathy would not be a high priority in my formulation of a response."
Go like his post (and learn more) if you also appreciated his response: https://www.linkedin.com/posts/kapilparakh_when-i-ask-people...
Perhaps it's rated as more empathetic because it's more likely to tell people what they want to hear since the patient is leading the questions and not the other way around.
That's another issue: since the A.I can't compute qualia; it can't discern if the patient has psychosomatic symptoms, so it is more likely to give a false diagnosis.
It gets a lot of things wrong - and I mean a lot of critically important things. Can a physician be wrong? Yes, but I can sue the shit out of them for malpractice too. Whom do I sue for bad "parrot advice" when GPT goes off the rails?
The only level 3 system in production (that I'm aware of, at least) is Mercedes', and they have liability while the car is driving itself. It shifts back to the driver 10 seconds (IIRC) after the system notifies the driver he/she must take over.
https://www.lipscomb.edu/news/lipscomb-co-creates-blockchain...
Sad statement on the judgment of the respondents. But an important reason it can turn out like this, I suppose, would also be that the RL feedback gives the model a fairly effective general optimization about what statements are liked by the mechanical turk-like evaluators. Most physicians have probably never had access to anything like that level of feedback on how their expressions are received. Maybe the LLM's can be rigged to provide goodness gradients for actual physicians' statements?
Nope, it's a sad reflection of the study construction.
Physicians' empathy was evaluated by their 52 word responses on Reddit. Unsurprisingly, a chatbot optimised for politeness and waffle outperformed responses of people volunteering answers in a different format optimised for brevity...