I used Claude Code to get a second opinion on my MRI
antoine.fi
antoine.fi
This really is key. We know we can't trust the AI, but at the same time we're also more comfortable asking the AI for clarifications or confronting it. Not having a time-bound appointment or paying by the hour helps a lot. But even then, more information doesn't necessarily help!
I once brought my 11-year-old car, a Civic with 150k miles, to multiple garages. I figured I'd play the "second opinion" game to correlate what the garages recommended to decide on what needed to be done...
I got 3 completely unrelated recommendations, including one that I knew was invalid! I felt worse off than when I started!
The solution to uncertain information isn't more information, which the AI can certainly provide, it's better information, and AI cannot currently provide that.
I'd argue that AI _can_ currently provide that, but that it can't do it _reliably_, and that to non-experts it's impossible to differentiate, which makes it all the more dangerous.
What is needed are studies that will take a cold look at the actual results because AI seems to be required to be perfect or it is useless. It just needs to be as good as a human for most stuff, but in the long run it will be much better. At least that what extrapolating current reality shows us.
Some of this might be applicable to LLMs, but some isn’t and much of it would be resisted. This is one reason we’re not likely to get “as good as a human” because at some level we’re not optimizing for the outcomes; we’re optimizing for speed, convenience, some participant’s economics, and underlying beliefs.
The LLM also gave me a bunch of questions to ask a new PT but I didn't have the understanding to judge the responses so I did more research. One of the things the LLM wanted me to ask about was questions about form and force closure and ideally would get a response about the oblique sling across the back. My PT didn't give me that exact response, but explained it in much more lay person terms, but because I had done my research I was able to validate their response was directionally correct. And so far, my experience with this PT has been much better. We're doing block pulls at 70% of my prior deadlift weight, next week we're going to go way back on weight and lower it some to get closer to proper deadlifting and work in some asymmetric loading exercises.
And in keeping with our theme of the limits of both LLMs and humans, unfortunately many exercise and medical professionals may not focus on how some specific genetic lots work until a catastrophic problem presents itself.
> ignoring feedback that a movement was causing pain, prescribing stretches to increase flexibility in a hypermobile person
So glad you found better advice than this!
LLMs are fantastic tools for exploring a topic, coming up with good questions and lines of understanding to pursue with professionals, and in picking professionals. I think they're also poor outright replacements for people especially when it comes to deep and important domains, but still quite useful.
Injury was to my SI joint, I've historically always irritated it lifting, but I set a new PR for deadlifts and it was debilitating for 3 weeks before my other half made me stop being stubborn and see a doctor about it.
I've also had my left shoulder joint surgically repaired after multiple dislocations.
When I ask a question outside of my domain of expertise I like to ask all of the LLMs I have access to. I also create separate sessions and ask the same question multiple ways.
It’s revealing to see how many different and contradictory answers I get, most of which are presented confidently.
The last time I ran a medical question through Claude I couldn’t even get consistent answers between sessions.
It’s also scary how easily you can lead each LLM to the answer you have in mind. When I would start asking questions about different options that other LLMs had presented, each session would drift toward that explanation.
You might end up with the answer from the most persuasive LLM, but you might also end up with better results.
Wonder if there is a paper out there on this.
I do something similar with reviewing code: I have one agent write the code and another reviews it, then they go back and forth for a bit improving the code. Seems to yield better results than one agent alone.
Seems like a similar principle.
https://www.nature.com/articles/s41746-026-02619-0
https://www.nature.com/articles/s44360-025-00007-8?fromPaywa...
Different prompt approaches and training doctors to use LLMs can improve accuracy of LLM-assisted diagnosis. It’s pretty reasonable to hypothesize that LLM “peer review” could improve that as well.
I never said it can't work. I just said that finding the correct medical digagnosis is different than finding a solution to a software problem.
Aside from the statelessness GP mentioned, one can insert anti-conciliatory intermediation. "I saw a random claim go by, but something about it seems not quite right. What am I missing? They said: [...]." Weaponizing the bias, and orchestrating the discourse from the harness.
It did great, generated a report on the assessed business that was incredibly detailed and plausible.
Then I started running tests and getting into the details, and found that if you ran the same report on the same data, it generated completely different, still very plausible, results. I could run the same source data through the assessment process 10 times and get 10 very different results. We had to can the project and go a different route.
LLMs are designed to produce plausible results, not factual results. We can fix this when using them for software dev by using linters and tests (though we've all had the experience where the LLM invents an API endpoint). I would not trust raw LLM output in any situation where that kind of testing and verification capability isn't present.
They are true to their name: Language models. It is precisely the same problem in a language: a grammatically correct sentence is not necessarily true.
When I ask an LLM, I trace the sources, and see if they make sense.
More often than not the sources don't actually say anything about the topic in particular...
> It’s also scary how easily you can lead each LLM to the answer you have in mind.
Exactly. Which is why "treat an LLM like a human expert who can answer your question" doesn't work. It's more like a human bullshitter who makes up convincing looking answers, and tries to please you. If the answers have actually some grounding in the training material, that's useful as some kind of holistic google, but often it's not.
Professional tip: you can cut out the LLM middleman here and save a lot of time and money.
It is very scary to me that people are entrusting potentially life-altering decisions to these things.
The problem with medical advice is that you may not be competent to verify the answer, right?
I agree that asking 5 LLMs to vote and trusting the answer is totally the wrong approach, of course. But LLMs (and traditional material) can help getting more informed. For instance, instead of going to your doctor with the LLM diagnosis and trying to convince the doctor that the LLM is right, you can try to build your own understanding of the problem and go ask the doctor to explain to you what you understood correctly and what you misunderstood.
If you have some understanding, it's harder for a specialist to bullshit you. But you need your own critical thinking and you need to put effort into actually learning something, blindly trusting and repeating what LLMs say doesn't help.
So the doctors gave him what he wanted: a treatment ... and Claude told him the treatment was a placebo. Correctly, I might add.
Yeah, it is absolutely not what the patient wanted to hear. BAD doctors! Except ... no, not really.
Does that explain what happened here?
> [the orthopedist] suggested I get an MRI, which the clinic conveniently had available. [...] This, of course, means little to me, but their suggested course of treatment was extensive; [...] Coming out of the clinic, I had the feeling they had jumped the gun.
The author also says:
> They injected me with Traumeel, which is registered in Germany as a homeopathic medicine "without a therapeutic indication".
I personally wouldn't want to be injected homeopathic medicine "without a therapeutic indication" without even knowing it is homeopathy.
And my recent experience with multiple doctors at multiple hospitals is that just like LLMs, you shouldn't blindly trust them. Sometimes they make mistakes, and in my experience they never, ever admit it (maybe even to themselves).
So trying to get informed "on the internet" (including with LLMs) feels sane to me, but that's worth what it is worth.
What concerns me in the article is that they have GPT make a diagnosis, Claude review it, and somehow seem to assume that if both LLMs agree, they must not be completely wrong. Just like for code, it takes an expert to leverage an LLM to take an expert decision. A beginner can leverage the LLM to understand the problem better, but they never reach the level of expert just from that.
Unfortunately it's part of lowering the confidence towards doctors, and the solution to that is to try and get informed and ask another doctor. And of course they don't like it if you say "I already asked someone else but I don't trust doctors in general, so I am now asking you to test the both of you".
This seems to be to be near the opposite of what a sycophant (ie. an LLM) would do.
This happens many times, and I usually have to lead the LLM through a chain of reasoning to prove to it that its objection, through generally sound, do not apply to my specific situation.
Someone not as well versed in the subject matter would think the LLM found a smoking gun (which they love to do), and be led on a wild goose chase.
What's the point of doing this?
But it does forget, and I'd have to prime it again for another session.
Scary in this context of course, but I find that it is an interesting thought for coding: it suggests that maybe, a developer who knows what they are doing will end up leading the LLM to coding something that make more sense than a developer who doesn't know and just vibe-codes blindly.
Sounds pretty obvious, but I wanted to say it.
A friend of mine's wife recently passed. They were chasing a suspected heart defect for over a year. She had been intermittently fainting. At about the year mark they decided to scope her digestive track. They found bleeding ulcers from cancer that was all over her body. I input her fainting symptoms into Claude and gastro impact was number two suspected after heart issues.
I have a few of other cases it's helped with. I'm not sure it could do worse than my own experience with the medical system. This is doubly true in places that lack any sort of medical care.
But also, I hear so many tales of running out of tokens. I ask Claude Code to build a tool to perform a task. I review the tool and then I let it rip if I'm happy with it. As I understand things, most just ask Claude Code to do the task. That seems a bit fraught.
Anyway, you have to impose constraints IMO and ask the right questions to get the answers you need or yes Claude Code (or any other LLM) will eventually just agree with you.
when i get an answer, and my first instinct is to ask a ton of follow-ups and "what about"s. i've learned to tamp this down with fellow humans, but with LLMs its great because most of the time the response is "you're right, something doesn't add up... let me try again". i think we eventually converge on to something reasonably true
So when asking difficult questions I tend to remove as much context as possible, rather than adding it. I don't want it to reflect my own ideas or biases back to me, I want an actually fresh perspective.
A mystery is worse. With each additional piece of data, the goal gets farther away. Everything is more and more confusing.
(Popularized by Malcom Gladwell)
Everything is a puzzle: there is one "Truth" or one diagnosis. You (a smart human) should be able to converge on it by cross-examining your LLMs. By themselves, they have no interest in revealing this, no stakes, which makes them tools only useful at the hands of a capable investigator.
What makes you think this is fundamentally different from cross-examining ELIZA? There is no guarantee that the LLM will help you converge on anything. Indeed actually calling out an LLM on BS tends to eventually produce an "I don't know and can't help you further" answer (as it should).
Absolutely. The guarantee does not come from the LLM. The LLM is a simply an improved version of Google Search.
The guarantee can only come from a systemic application of epistemic discipline and reasoning, which is very much (smart) human territory.
Put it another way, I could make good decisions with/without LLMs, with some uncertain diagnostics as input. I would have to trawl through 50 papers myself, and it is possible that my decision arrives 5 years too late as a result. LLMs enable trawling and do some of the legwork in connecting the dots, but are ultimately only as capable as the orchestrating human.
This is the primary business model of enterprise IT and is why companies pay so much for 4 hour disk replacement.
I get it - getting an opinion from a mechanic is time consuming. Not true of AI though.
He said it better using an expression I hadn’t heard before or since, something like “don’t go looking for goats when your herd is already with you.”
I read about before there's proper engineering / physics theory about this too, it's like a car as a machine is a linear/smooth physics system with multiple weaknesses. Overtime longtime period of running many places might weaken but it still evolves into a slightly different smooth system, until you introduce a replacement which cause a mis-match of impedance or something like that.
You’ll do something to prevent a failure (like, replace an old but functional alternator) but cause an oil leak or engine vibrations because you had to remove the propeller to complete the job.
I almost had a very similar experience with my beater Lexus. It took 2 independent shops and 3 dealers to finally figure out what was causing the ABS to go off randomly at low speeds. Turns out there's some obscure Toyota-specific tool from the late '90s that picked up a proprietary diagnostic code, and the third dealer was the only one that still had that particular piece of equipment.
...and of course, the thing that's broken has been out of production for 20 years and remanufactured ones cost more than the car is worth. I ended up just unplugging the ABS control module.
Point being: once I knew what was wrong, all the seemingly contradictory information from the other 4 shops suddenly fit together. It's just such a weird thing to go wrong that no reasonable tech would ever have considered it.
It sometimes can, if it straight out never can no one would use it. People use it , lots of them.
The AI might be very good at diagnosing all minor issues, but might not lead to a successful repair, whereas human mechanics are extremely good on 80% of major issues that's not the ground truth, but will lead to successful repairs (that might not address the root but simply patch it). So it comes down to manage expectation / outcomes.
I would frame it differently: you now know which shops are not to be trusted. So, next time you need one, you will take a better decision.
Aside from the LLM-ism (it isn't foo, it's bar) - this is a thought terminating cliche. You definitionally don't know if some information is better or not given that you were uncertain about the information in the first case.
"I went to three mechanics and got three different answers" - your takeaway is just "Ah - I clearly need better informed mechanics."
Which is on it's face absurd because if you could clearly judge the ability of the mechanics you wouldn't need their evaluation. You'd just do the evaluation yourself.
Scammers who do the lowest effort diagnostic and "fix" to get you to pay a smaller amount of money to fix the problem in the short term even though it'll re-present itself a week/month/year later.
Upsellers who will find other things "wrong" with your car and pressure you into paying to fix them because they sound a lot worse than they are.
Good mechanics that will explain what they did to diagnose the issue and recommend different options depending on what the issue is.
Funnily enough, I've found that doctors tend to also fit into these 3 archetypes.
I have read about doctors complaining that "with AI, patients now come with their own diagnosis and don't trust us when we say it's bullshit, and it is a problem". I can feel for them, but if they give the feeling that they don't listen to the patients and the patients don't trust them, it's not only the patients' fault, I would say.
[1]: I have more than one examples of my relatives like this: A doctor says "wow that's bad go to the ER", the ER says "nope it's all good, go home", first doctor learns about that and says "WTF you GO TO THE ER, call me and I will insult them on the phone", and finally resulting in a surgery where the doctors say "they were lucky we could operate right now, because in a matter of hours they could have died from this". How in the world can I trust them after one event like this? Happened to me (in some variation) 3 times. Not based on an LLM diagnosis in the first place: based on a doctor's diagnosis.
For example, I just prompted Kimi-K2.6 with:
> I'm considering buying a used base model 2010 Honda Civic with 80k miles that's been garage kept. What are typical things that go wrong with this type of vehicle?
It listed 10 issues including the engine block cracking (which wasn't even an issue with 2010 Civics). Started a new chat and asked about a 2010 Toyota Camry, another unbelievably reliable car, and it listed 9 similar issues. Started a new chat and asked about a 2011 Jeep Grand Cherokee, a notoriously unreliable vehicle, and it listed the same number of issues.
Sure it's data to make decisions on either way, but it really all comes down to how good your prompts are and whether or not you can think critically about the output, whether or not that output is a conclusion or just data collection.
Before I was admitted, I quickly found another radiologist, who diagnosed pneumonia instead. I sent his report to the chief doctor at the tuberculosis hospital, and after some deliberation they concluded that the original reading was wrong. Turns out the doctors there can't read scans at all and just believe whatever a radiologist says...
The funny thing is, they had already officially put me on the tuberculosis register and didn't want to admit they had made a mistake. So instead, they simply gave me another paper saying that I had been cured of tuberculosis by them... in 7 days. I'm probably the only person in the country to defeat tuberculosis in a week :)
So if you don't trust the radiologist/doctor, maybe find another doctor if you can afford it? You can compare their conclusions and see if they match. Two unrelated doctors or radiologists saying the same thing is probably about as close to the truth as you're going to get. I'm not sure though whether I should trust AI or humans more. AI can hallucinate, but I've been misdiagnosed by humans so many times too...
I forgot to mention that, besides getting a second opinion from another radiologist, I also took a more modern test at another private clinic. That test has better detection rates than the one the state clinic used, and it came back negative too.
I have suspicions they had some kind of government quota to keep the hospital staffed with patients in order to receive funding. Or they were just completely incompetent. I pushed back by bringing them another radiologist's report and the results of a better test that I paid for myself, so I guess they decided to back down.
Think about the consequences of mistakes in both directions ...
I suppose there could be reasons, but I don't know them.
https://theunion.org/news/is-involuntary-incarceration-of-tb...
>17% said that, as a matter of principle, the involuntary incarceration of TB patients was inappropriate on any grounds.
>Regionally, members from Europe Region had the highest percentage of respondents objecting to the policy as a matter of principle (26.2%) while the North America Region had the lowest (3%).
The emergence of multi-drug resistant tuberculosis in the 1990s is probably one of the reasons:
>Respondents most strongly supported the policy of incarceration for patients known to have multidrug-resistant TB (49.7%)
But AI as of right now is worse than any bad doctor I've ever worked with.
Anecdotally, several people in my life who embrace less traditional (and sometimes more conspiratorial) views on modern healthcare tend to be the ones that can't afford it. A confident-sounding chatbot to answer questions day and night about what's going on with your body is very seductive in a world where access to real healthcare is getting further and further out of reach.
Everyone is either a "all doctors are scams" QAnon type, or they blindly trust everything their doctor says, no matter how fishy, in fear of coming off as one of the former group.
And, to use a phrase we all hate by now, you're absolutely right. When most people have to go into debt to even see a doctor, what can people possibly conclude from that besides "all doctors are out to scam you?"
They have at-home COVID+Flu tests are my local CVS for $35, why go to an urgent care?
Anyway, thanks for sharing!
I've heard this experience from quite a few folks before, but this is my first time hearing about a mandatory 8 months quarantine as a consequence... damn
Just FYI that might suck in the future. Some countries might deny you entry / visa, or at least require expensive additional screening if you had tuberculosis diagnosed in the past. It might be worth to fight this.
These models are generally terrible at reading medical images. The amount of public training data on the internet compared to the number of scans a radiologist reads in training is minuscule. There’s obviously a ton of medical images in general but very few, and even fewer along with a report are available on the internet publicly for download.
There are vision language models coming out of research labs that are excellent in describing and localizing findings. Still at the level of a 1st or 2nd year radiology resident, but as we all say - this is the worst the models will ever be.
General-purpose LLMs are _fantastic_ at medical diagnosis that do not involve imaging. I am completely convinced that given enough information and time, frontier models already outperform >90% of doctors on initial diagnosis of internal issues and suggesting medical tests to further reject or confirm the most likely theories. To the point where I'm eagerly waiting for the first hospital in the world that's willing to be open and honest about using them for that first step, and then proceeding from there. I'll be on a flight there as soon as one arrives.
At the same time, they're worse than useless at anything involving medical imaging. Asking them to interpret them is worse than trying to interpret them yourself as a layman. And you surely wouldn't interpret them yourself.
> General-purpose LLMs are _fantastic_ at medical diagnosis that do not involve imaging.
Can you share the reasons that you believe this? > At the same time, they're worse than useless at anything involving medical imaging.
What is special about medical imaging that makes AI/LLMs specifically bad?It's multiple things. It never shows the subscapularis in the way that people actually look the tendon. It hyper fixates on the axial when I find the sagittal much more useful for subscapularis.
Figure 7. There's an arrow pointing "to the acromial undersurface". The arrow is not pointed to that location.
Figure 5. "thin bursal fluid". This is within physiologic variation, but is calling bursitis.
It keeps bringing up irrelevant normal things like the shape of the coracromipal arch, I assume because lots of websites have information about that as a patient focused possible cause for rotator cuff impingement.
I am reminded of the recent Stanford MIRAGE study which found that LLMs will happily hallucinate answers about medical images if the medical images are omitted.
Firstly, please keep in mind I'm talking about the entire doctor population of the world here. Not sure which particularly bubble of this earth you have experience with, but note how half the word's population lives in India/China/Indonesia/Pakistan/Nigeria/Brazil/Bangladesh/Russia. Now I do believe that it holds the same for e.g. Europe and non-China East-Asia, but still.
How many patients has the world-wide average doctor seen? How long have they been a doctor?
How many have they seen with the particular condition the patient has?
How much time do they spend listening to and reasoning about a patient? The median in the world is likely under 3 minutes.
How many real-world incentives do human doctors have to deal with?
Given infinite time and resources, and zero external incentives, maybe the median human doctor would outperform the LLM at this task. But this is completely detached from the real world.
> What is special about medical imaging that makes AI/LLMs specifically bad?
LLMs: Besides lack of training data as mentioned elsewhere, they're simply not trained for high-fidelity image processing in general. It's not limited to medical imaging. It's a bit like the "How many Rs in strawberry" thing, but worse.
As for "AI" in general, medical image analysis is a very active field. These tend to be purpose-built though, not general-purpose. It seems likely at some point they'll become mainstream, but there's still a way to go.
There are now open source open weight “foundation models” coming out of labs that are transformer or mamba based architectures. These will accelerate development.
Like OP, I also had a shoulder MRI, and asked two AIs for opinion (awaiting a follow up appointment to discuss the results).
They both insinuated much more serious problem than it was (as judged by an orthopaedic doctor).
Even for diagnostic, I’m not totally convinced AI is going to cause a lot of issues. From a medical-legal perspective I doubt AI companies are willing to take on the risk of misdiagnosis . When that happens, human rads will start to get phased out.
In the interim, I think the next 10 years could be a golden period for diagnostic rads. Rads will still be the ones doing the work and signing reports, but people who learn how to use the right combination of tools will become very productive and can make a great living. Eventually payers will re-align but early adopters who figure it out will have a leg up.
My understanding is that medical images are part of a patients record (so they must be available for the patient or other docs at the request of the patient) but whoever has collected the images does have some form of ownership. I’m not a lawyer and I have a cursory understanding of this. I believe it would be possible for an AI company to lease or get access to the data through a transaction but I suspect that it hasn’t happened (or happened publicly) due to fear of backlash.
For example I know that some companies like Tempus had access to imaging that corresponded to tumors which had been biopsies for sequencing and they were developing models in house.
> They performed shockwave therapy on my shoulder even though a recent clinical practice guideline says clinicians should not use or recommend shockwave therapy for rotator-cuff tendinopathy without calcification; I was told during ultrasound that there was no calcification.
Ultrasound isn't a great way to assess for calcification. It'll find large calcification but easily miss small ones. Plain radiograph would be more helpful, but the MRI may have revealed it as well. Either way, shockwave therapy isn't harmful in the absence of calcification--it's just not helpful.
Edit: when a radiology report says something isn't present, there's always an implicit caveat that the finding isn't present within the context of the modality and images obtained. So an ultrasound report can state there are no calcifications while a plain radiograph can report the presence of calcifications without being inconsistent. Obviously very confusing to patients and people unfamiliar with medical jargon, but clarifying this in reports would make them sound even more qualified, "hedgey", and annoying to read than they already are.
Edit: I should mention that ultrasound is basically unusable for evaluating bones. Sound waves can't penetrate bone, and so you end up just seeing a huge black void. That's a huge orthopedics use case that ultrasound just can't benefit. However, ultrasound is fantastic for evaluating muscles, ligaments, tendons, and other superficial soft tissues.
Since MRIs are more expensive, private doctor's might order them instead of an ultrasounds.
(I'm a doctor)
Given the skill involved, it's probably a liability concern they don't want the exposure over there.
Spoiler: because it hurts like hell.
Both inductive and abductive reasoning would say that just because something hurts that doesn't mean that everything that hurts is that thing.
Any comment that doesn't start with this or similar qulaification should be taken with a grain of salt (yes, including this one).
Medical imaging is one of those things everyone thinks is simple because they don't know what they don't know. I'm a cardiac sonographer, and I have to assume radiologists hear at least as many eye-rolling takes on AI coming for their job as I do.
Full sarcasm, is there one that’s that’s more immune?
I know my anatomy and etc and have done a short stint in ultrasound. I have no idea what you are doing or looking at and can identify pretty much nothing.
Echo techs are going to be around a lot longer than MR techs.
This is being overly nice, I think. Anyone who doesn't understand this is an idiot imo. You would have to assume that every type of diagnosis instrument has infinite clarity and is always correct to be confused in this case.
Reminds me of the Babbage quote where somebody asked him, if I put the wrong question into this computing device, will it still give me the right answer? His response, paraphrased "I can not fathom the logic of the minds which would come up with such a question".
Would have also been a fair point if Babbage had channeled his inner techbro and insisted it would directly replace human calculators; simple machines like Babbage's will chug along blindly on obviously erroneous data, but humans for all their sloppiness can often backtrack on errors.
Well, he did diagnose the situation correctly. He couldn't comprehend the confusion of ideas that provoked the question.
I'm also not entirely sure it's an odd question to ask. To this day, users are surprised when their software produces garbage output instead of failing. Perhaps the members of parliament were expecting some form of input validation or sanity checking out output.
I suspect their sarcasm might have escaped Babbage who seems to have been on what we now call "the spectrum."
Isn’t there a saying about there being no stupid questions, only stupid answers or something?
I disagree. A priori it's not obvious to a layperson whether or not a statement that uses unconditional phrasing is intended to be authoritative or conditional on something unspecified, like the resolution of the measuring device. This goes for any sufficiently technical field.
If you got the brakes checked on your car, and the mechanic did <something> and told you there are no issues with them, and you then took your car to a different mechanic who did <something else> and told you there is a problem, you would not be an idiot for thinking that these conclusions contradict one another.
I don’t think that’s true. Avoiding this mistake requires knowing that an ultrasound may not detect calcification. For a patient reading their own report, I don’t think that’s intuitive. I would expect most people to read “no calcifications” and assume that their joint has no calcifications.
If a report states that X was not found, it does not mean X did not exist, it means it was not found.
What may be lost on the layperson is the nuance and understanding of how thorough or not a particular scan is and how much weight to give the findings and thus the odds that the report is correct.
Yah you can argue that the tool is not ideal for that diagnostic, yadda yadda. I get it, and in the end I agree with the subtle difference you highlight, because it is something that makes sense to a certain kind of people. You know how many medics would read the report exactly like the author did? Too many.
How do I know? Im not in a wheelchair after being constantly misdiagnosed by using the wrong imagiology technique by (mostly) chance, and a good help from friends, including a surgeon. This seems to be a case where AI would be a valuable doctor tool for differential diagnosis; instead we have know-it-alls that can't bother to verify, and AI that often gets details wrong. That is the problem.
That might be true, but it is definitely not the world we live in.
And I think we can easily have examples where we can reasonably trust this, and a spectrum of such.
E.g. there is a math solution and the report says "there is no errors in this solution", you would imagine that to be quite reliable, no?
Even then, the context that "ultrasound isn't a great way to assess for calcification" is important when reading either statement. Laypeople don't necessarily have that context.
I’m fairly sure that there are no lions in my house. Lions are quite large and I’m capable of detecting lion sized objects with my eyes.
To demonstrate that something is not present you first define the object, then come up with a test that will reliably detect the object. If the test comes back negative then the object is not there.
In a strict philosophical sense I cannot prove that there are no lions in my house, the external world might not exist! A hypothesis that no one has thought of might be correct and that hypothesis could show that there are invisible lions in my house!
However I intend to act with the certainty that there are no lions in my house. Because I have no evidence of lions in my house.
Absence of evidence is evidence of absence.
My internal model is/was “if the scan wasn’t set up / can’t detect the thing, why would the statement be present at all?”.
That implicit assumption is really subtle.
There's a difference between 99.9% clarity and 50% clarity. Even if neither exactly equals 100%, it's understandable that a layperson would expect different language between them
Even if this is true, so what?
Idiots get sick at least as often as others, and the medical system needs to work as well as it can for that population too.
> seriously underdeveloped
What's the difference?
Someone on reddit claiming to be a radiologist claimed that.
I wonder where the savings will go when those jobs are gone.
The radiologist I know does not, but they are paid very well (and these numbers are always dumb when you're not sure if they're living in Manhattan vs literally anywhere in Kentucky)
Like most medicine, a large % of the job could be done by any decently talented person willing to follow instructions and shadow for a few months.
Like most medicine, the remaining % is what you're paying for, because it is literally life and death and you can't do things like "pull the logs" or "lets turn it off and take it apart" or "huh i need to put this down and come back later". Even in radiology, because "well lets just do it again to be sure" is often not a viable option.
While there is a problem in how we have inflated the cost of education for medical fields, the insane health insurance issues (US obviously, but it does have some effect globally when the expert radiologist you hire from the US to help with research costs that much), and probably some better ways to approach splitting the work for the entire field, like most professions dealing in life or death, medicine likely will always be paid well.
You would think that the AI would point out that calcium is best demonstrated on Radiographs/CT imaging vs Ultrasound or something to that effect.
Between the toes and the below the knee amputation, there were no less than 15 different doctors and PAs / related personnel who COULD NOT COME TO A CONSENSUS. They would just tell my mother and I (PoA) the details; they refused to come up with a singular plan of action moving forward, leaving it up to us to make 'an informed decision,' something that's IMPOSSIBLE when you have to take up to 15 different opinions into consideration.
What exactly are we supposed to do as patients/family members when medical personnel cannot give reasonable paths forward and instead just throw a bunch of shit over the fence at you and tell you, "you decide what to do from here," regardless of how many VERY DIRECT conversations I had w/the 'care team' on doing better to provide a limited array of options and reasons/likelihood of 'positive outcomes'.
I'm used to dealing with a wide variety of stakeholders/SMEs in decision-making; it's my job to apply my extensive industry experience to present our clients with their options, ranked and reasoned. Doctors, in my experience and most recently with my father, clearly do NOT do that (I assume due to liability; but, no real idea, honestly). So; when dealing with LIFE CHANGING circumstances, what are we supposed to do except rely on what might be able to offer more analysis and option narrowing w/AI?
I certainly don't want to make the job of medical staff more difficult by putting out crazy theories I found on the interwebbernets through my own research, etc; but, when we're having to deal with uncertainty and insanity, what else can we do?
All of the above was intertwined with brief stints with doctors that would just berate her for being a painkiller junkie, even though she hated the stuff and just wanted to find/fix the problem.
Kind of a rant, really. I'm not sure how to tie it back into AI. I do wish we had AI at the time so that we could at least cross-check, but I also understand that doctors are already sick of patients self-diagnosing on the web and that AI probably just makes that worse. At the same time, if our medical system could catch up a bit (more doctors? less corruption/paperwork? not sure what it needs) then maybe people would be less inclined to take matters into their own hands.
AI is absolutely a god send for patients navigating the medical system.
I know the US system is horrible and I sympathize with doctors doing their best within it. But we must admit, they are also responsible for the countless stories just like yours, and have contributed to the public's deteriorating trust of medical institutions. It's not just the insurance companies and conglomerate CEOs.
P. S. Opiates work, very well. If you tale one and it doesn't work, you might also be allergic. Hydro didnt really work, then itching and puking started after 2 days of untouched pain(essentially, a reduction of 2 points). Codeine nausea.
I wasn't aware of the possibility.
I empathize with you as my grandmother is also being what feels like gaslit by several doctors being told her symptoms are dementia and not from the chronic UTI’s she suffers from, but when the UTI’s clear all of a sudden no symptoms. Our medical system is very frustrating, between doctors who don’t have the time needed for complicated cases, or the threat of every patient suing them causing them to be overly cautious.
I think as long as you’re aware of the pitfalls of AI, as you seem to be, it’s a solid tool for helping to understand medical situations with the right amount of double checking.
I don’t think our system will improve until 1) We increase the amount of doctors in our country and tell Congress to quit limiting the amount we can train yearly, and 2) the system of liability for doctors is changed to be more like New Zealand’s where their liability insurance is nationalized and they’re not at constant threat of losing everything to a lawsuit over a patient who wouldn’t take their meds and got worse (the cases are generally much more complex than that, but the idea is it’s not always the doctors fault as media would have us believe).
The problem is these are just statistical models at the end of the day, so you need to know something to be able to identify the errors. You can’t let them really be autonomous and you also can’t really have people turn into glorified approvers. If the machine is correct 89% of the time, you cannot make people responsible for that 11%. It’ll just cause automation fatigue.
tl;dr: the actual use cases of these LLM (or generative AI in general) is rather limited, so it is offensive how much hay has been given to them eating the entire capitalist system. They are not fit for purpose.
The human experts are literally just a trained biological neural network. In this domain they are not capable of anything a computer can't already do.
Humans can identify. A computer vision model can return a statistical value. Both can make errors, but these errors are orthogonal to how we work and what is being asked of them. I think a CV model can absolutely provide value as augmentation. Identifying possible misses or a different diagnosis worth considering, but that is not what is being asked of them here. The pitch by Altman and Amodei is not to say, “This tool that might cost $1,000/month can help increase the accuracy of your diagnoses by 10%,” instead it’s, “This tool can allow you to keep 10% of your workers to monitor it and you can fire the rest. Also, the workers carry all the liability.”
> The human experts are literally just a trained biological neural network. In this domain they are not capable of anything a computer can't already do.
People need to stop making this baseless claim. Human beings are not stochastic computing devices, we are not neural networks. We don’t fully understand human cognition and intelligence. I have the highest confidence we will figure it out one day, though.
Yes, neural networks were based on a superficial view of the human brain, that’s it. For instance, it is biological impossible for the human brain to do backpropagation, which is kind of important for a modern neural net.
This really rubs me the wrong way because it's objectively false, but people keep bring it up because I think people want it to be true rather than accepting generative AI for what it is: a tool with a bunch of caveats.
Also disagreement among human radiologists has been documented for decades, so the clean expert baseline you're defending doesn't actually exist outside this argument.
When the identifiers pass human-detection-rate percentages, it will most likely be cheaper to hire a fall-guy for the liability with a much smaller salary, I think this will be a big market in fact.
I also didn't say anything about whatever Altman or any specific company is doing.
The simple fact is that we send humans to school for years to learn to read and classify these things. It's something computers will be able to do strictly better.
I don’t think there’s any evidence that’s true.
The premise seems to be proven, just not extended to the medical domain yet.
Satellite image processing can certainly detect a hotspot then interpret it as either a small brushfire or a missile launch. Facial recognition detects my features then interprets who I am. It's all pattern matching just at different scales in different parameter spaces.
If we would allow AIs to be trained on the petabytes of medical data hidden in hospital systems, they would most likely be much better at diagnosing illnesses and conditions than the average doctor.
(Justifiable) Privacy around medical records so far prevents this.
You think you're cheering for humans, but in fact you are gatekeeping healthcare.
It would be like the sum of all medical professors in existence.
Given human body complexity, the diagnosis is a compound output of the experience, knowledge gained throughout the career and diagnosis methods/equipment, the title (like Dr) is a certification imposed by the state so its "safe" to let people practice since they passed "the bar" - but that doesn't imply everyone will be treating the same.
Some specialists update their knowledge monthly, some yearly and some don't do it at all, there are so many variables in play here (geo, politics, even weather haha).
Having said that, choosing the specialist is really important, getting opinions about their practice and their speciality, you can only maximize your chance of getting the right diagnosis, but don't expect to get it right just because somebody is called a Dr.
There is absolutely one "The Diagnosis". Human body is a machine, albeit a very complex one, and all measurement sources have noise. But they are all measuring one reality, and if there is a problem, there should be one explanation that all measurements align with. They can be noisy but can never be conflicting (instrument error notwithstanding).
Physicians' ability to arrive at "The Diagnosis" would vary, but it does not mean one does not exist. I am not sure if characterizing human body as derministic or not is relevant here.
Thus, chasing the „right” diagnosis (whatever that is?) is pointless, as it only the outcome (reducing symptoms, stopping the damage) can tell you if the diagnosis was right, but not the only one right.
"The Diagnosis" does not mean "one root cause".
Situation: my car has some unexplained vibrations. 1. Mechanic A says that it is the engine mounts 2. Mechanic B says that it is some weirdness in how the exhaust assembly is hanging to the underbody 3. Mechanic C says that it is just my wife farting
I replace engine mounts and 40% of the problem is reduced. I then drive without my wife and the remaining 60% is solved.
"The Diagnosis" was: 40% mounts, 60% wife, 0% exhaust.
There is always one "The Diagnosis".
No, that is not true at all.
This is a kind of thinking a lot of programmers fall prey to. The real world, outside of code, is a very fuzzy and inherently analog place. There is very rarely one in any complex system having a complex problem needing a complex solution. At some point even the definition of diagnosis gets fuzzy.
The best demonstration of this in medicine is probably the DSM-5. What, really, is the difference between Narcissistic Personality Disorder and Borderline Personality Disorder and Generalized Anxiety Disorder? Can they overlap? (Yes.) How do you treat them? (It's not easy.) What about depression: how do you tell if someone has Major Depressive Disorder or Bipolar Depression? (Again: not easy.) In some circumstances the only way to tell the difference between the two is what drugs work: if antidepressants help, it's Major Depression; if mood stabilizers help, it's Bipolar Depression. It's kind of odd to define a One True Diagnosis by "well we fixed it this way, so it must have been that", with no other way to do it, isn't it? (What if both work? What if one works for a while, then the other works? What if treatment with antidepressants induces bipolar (hypo)mania? All of those happen!)
And that's just a few examples.
> This is a kind of thinking a lot of programmers fall prey to. The real world, outside of code, is a very fuzzy and inherently analog place.
Having said that, I would vehemently reject and push back against this, and without doubting your sincerity, characterize it as an ad hominem.
The vast majority of issues with the human body are mechanical in nature. Restricted blood flow, unwanted tissue, a broken bone, a bad valve etc. These are causal descriptions of "disease". Where causal descriptions exist, the "One True Diagnosis" principle holds. Psychiatry just happens to be unique in that it is a fuzzy science where we rely on checklists and ultimately all diagnosis is probabilistic.
EDIT:
> This is a kind of thinking a lot of programmers fall prey to. The real world, outside of code, is a very fuzzy and inherently analog place. There is very rarely one in any complex system having a complex problem needing a complex solution. At some point even the definition of diagnosis gets fuzzy.
I would also push back against this mindset in general. This is not a falsifiable claim, it is incoherence as an argument, and I do not need to be a programmer to hold this position.
That the real world is analog is irrelevant to its amenability to causal explanations. Or "fuzzy": "fuzzy" in this context just does not mean anything.
I am not trying to sound exasperated or win internet points, just impress this point on you and anyone reading this. We can write math to predict weather, make it tractable to solve using approximations, tolerate IEEE 754 weirdness, and finally tell what the clouds will do a week from now. This is nature telling us that there is a pattern to how it behaves, and it is the only weapon we have as scientists.
To say that nature is not amenable to explanations is a very defeatist thing to say: neither Newton nor Einstein nor any of the million-odd people that have built modern society would exist if nature did not have causal explanations. I urge you to reject this defeatist thinking.
Ok let us unpack this statement.
For your point to hold, I would have to be saying "all kinds of practical diagnostics are invented now. No progress can be made in better diagnostics".
If Alzheimer's can be validated by slicing open a dead patient, there is a causal mechanical explanation for the disease. If we can not confirm that defect without slicing open the patient, that is a limitation of 2026 tools. The "One True Diagnosis" is an Oracle explanation that all real diagnostic techniques try to approach in the asymptotic sense, and it is helpful exactly because it clarifies in discussions like this.
There are going to be diseases where we do not yet have causal explanations. Or where we treat them without establishing them. Hypertension is one example: while technically it can be caused by vascular stiffness, some weirdness with the RAAS system, some hyperadrenergic weirdness, practically you get a lot of mileage out of just prescribing people telmisartan if they're old.
That does not mean the frontier of hypertension is settled, or the 10% who do not have a vascular stiffness problem would not benefit from better causal models of hypertension. Science is us continuously pushing back against the fog: of the tools we have in 2026, some are great, some are imperfect, some are promising etc.
Regardless, to bring the discussion back to the claim at hand: at all points in future, we will need the ability to reason under partial information. "Absolutely flawlessly complete diagnostics" is an asymptotic goal we get closer to but never reach. This is both very doable for a disciplined human, and very hard to outsource completely to an LLM. Treated as tools operatored by competent users, they are magical. But they can not outperform their user.
Even so, we’re operating on approximate datasets and sometimes our predictions are wrong. I think a lot of the medical field is like that - people are doing the best they can with what they have.
It’s entirely possible that DSM-5 will be viewed as flawed and inaccurate in a century, but it’s better than nothing.
Similarly, for every possible medical affliction there could be “The Diagnosis” that would describe how to treat it, we’re just unable to be that accurate and thorough. The fuzziness just means that you’d need 10’000 data points about the state of the body instead of 10-100 and also be able to reason about them.
Quantum mechanics is an excellent example here. It is not "defeatist" to accept that we don't know where the electron in the hydrogen atom actually is. Or to accept that if we really, really wanted to figure out where it is, we can only do so by disturbing it enough that its position is probably no longer a useful thing to know.
These are fundamental features of the world (at least according to our best theories of Nature), and it is only by accepting them and the uncertainties inherent to them that we are able to make progress using those theories. Among the consequences is that thinking about "the position of the electron" is not so useful; we instead need to leave position behind and start thinking using a new thing, "the orbital of the electron". This is a major conceptual change, and internalizing it can be Very Difficult for some people.
But the world does not care. It and its complexity owes nothing to anyone. It is us who must adapt to the world, in all its fuzziness and incompleteness. Nature will break the rigid, but if you bend, you can soar.
> In some circumstances the only way to tell the difference between the two is what drugs work: if antidepressants help, it's Major Depression; if mood stabilizers help, it's Bipolar Depression.
This is ridiculous. There is zero mention in the DSM-5 or ICD-11 of "if these drugs work, it's this, otherwise it's this." I would question a psychiatrist dispositively making a diagnosis on such grounds.
In a community largely made of people whose job it is to produce such functions, I'd say it's to be expected
There's no shortage of tech people convinced they deeply understand law, medicine, philosophy, etc. despite never having read much on the topics.
I can't find it but one of the greatest show HN was a blog post about someone who was annoyed by his inconsistent shower temperature control. From memory, he spent a full weekend adjusting it, taking measurements, making graphs, and proposed "next steps" about prototyping better temperature control with microcontrollers and servo and pontificated about developing a product, of course controlled by software. He skipped the part where a bit of research leads you to the already common "thermostatic mixing valve".
I also had a pretty painful shoulder issue at one point, where the pain just wasn't subsiding for months. I tried massages and acupuncture as I didn't want to do surgery, but it wasn't helping at all. The thing that fixed it for me was just really focusing on doing pull-ups. I couldn't do them at all when I started, so I began with dead hangs and scapular pull-ups, eventually progressing to regular pull-ups, and then training with a "grease-the-groove" method once I could get a few per set. I stopped the training schedule once I was getting in around 17 pull-ups per set, and now just do 6 sets of about 7-8 pullups 3x per week spaced throughout the day. I'll also do some shoulder mobility drills [1].
Whenever I get lazy about keeping up with them inevitably discomfort will start arising again, but it goes away once I get back to strengthening.
It really seems like if you, as a patient, go looking for a quick fix, that’s what you’ll be offered. And if you educate yourself a bit and then go t for the best fix for you, you usually get they.
With calcifications, physio without the shockwave component definitely doesn't allow going back to the normal gym routine. It's just not enough.
Strengthening with PT kept the joint stable enough to stop rubbing and allow the inflammation to clear.
And as long as I stick to a regular gym routine that includes rotator cuff work, it doesn’t recur (and did the few times I lapsed).
But absolutely, PT doesn’t fix everything. Bit for a lot of things, it’s worth trying - but it might also means a lifetime of altered habits to keep whatever injury/problem from recurring.
I broadly agree though; about a decade ago I had the standard office worker low back pain problems which cleared right up after doing squats multiple times a week. Of course a decade later I managed to blow out a disc at the gym, which I still work through as I write this today, but well worth the risk in the long run. Even with that long experience of strength training, the PT was worth it even if it didn’t fix my problem entirely. It added some variety and pointed out some details I had overlooked to improve my shoulder health.
In my limited experience, "If all you have is a hammer, everything looks like a nail", rings particularly true with medical professionals.
in the end I went with the surgery and now I'm in recovery. Sadly 4-5 years of trying lots of PT/streaching/pull-ups/push-ups/whatnot didn't help and only made things worse :(
The respectable ones know they aren't doctors, but they've seen a lot more recoveries and cases where minimal intervention was required. As some people have said some surgeons like to cut people up.
I have hip impingement, I played sports and got a labral tear in one of my hips. My hip would get sore and painful after a lot of activity. I've seen a top surgeon in the US. After we just met, he looked at the MRI(yes there was a small labral tear there) and said he can have me on an operating table in 2 weeks.
I was shocked, because the recovery absolutely sucks. So, I got 2nd and 3rd opinion.
3rd opinion was a doc with 20+ years of experience. Asked me if I plan on going pro in any sports, I said no, he said the surgery is not worth it. I did some PT and barely have issues with that hip.
Then, Obama admin created a website to see what gifts($) physicians accept. The 1st surgeon had accepted six figures+ from stryker. The older doc? 0.
There is no money in PT for a surgeon. I would thread lightly with popular and young surgeons.
ChatGPT surfaced a NIH study that concluded that 20% of people have allergic reactions that are isolated to a body location, and that shoulder "skin prick" testing may not reveal. I asked him about that and he said "that's not how allergies work". Full stop. He was unwilling to even look at the study.
He prescribed a CPAP and regular nebulizer treatments. Side story: the CPAP place sent me a SMS message that I couldn't recognize was not a phishing attempt, and when I reached out to inquire who they were they never replied.
So I decided: Let me just try taking a second-gen allergy tablet every day and see what happens.
My sinus infections have gone away. Previously I was getting a major sinus infection at least quarterly. Maybe he's right that allergies don't work that way, but allergy tablets have absolutely solved my problem. Which I'm thankful for because I tried a CPAP for a solid month a few years ago and I just could not get used to it, and was sleeping like crap.
All I can find is about 1st gen antihistamines (i.e. Benadryl, which I doubt many people take daily, because of the drowsiness).
Even for those, evidence seems to be mixed at best. "Huge increases" seems like hyperbole.
I think this post is a decent summary, the answer is a soft maybe: https://www.health.harvard.edu/mind-and-mood/should-i-worry-...
Second-gen tablets might increase dementia risk by a small-to-medium amount (there's almost certainly still a small degree of CNS activity, and we don't know what causes dementia in the first place), but researching it will be difficult. Dementia is hard to research because of how long it takes to develop, and it's poor coding within health data, and antihistamines are hard to research because they're not often prescribed and aren't available in the health data.
If it's a large effect, those factors wouldn't matter, but smaller risks are harder to detect and more sensitive to bias. If you want to minimise dementia risk, then reducing antihistamine use might be warranted, but you're probably better off addressing the risk factors we do know about: https://www.dementia.org.au/brain-health/risk-factors-develo...
You didn't read it very closely, I posted it specifically BECAUSE it cites sources (14 of them by my count).
edit: I sit corrected, if I open it in Incognito mode the citations are removed. That's not very useful.
edit 2: Here is a gist with the citations: https://gist.github.com/linsomniac/6d2bdeb0f63cf504354b067e2...
Only first-generation antihistamines with anticholinergic effects are associated with cognitive decline in elderly patients.
Yet here we are, warning each other about the dangers of LLM hallucinations. Humans "hallucinate" (provide random authoritative-looking information without anything to back it up) pretty often too.
https://www.myalzteam.com/resources/zyrtec-and-alzheimers-me...
There IS one year-old finding that suddenly stopping Zyrtec after daily 3-month use may lead to nasty itching, and if that happens you can re-start and then taper off. https://www.fda.gov/drugs/drug-safety-communications/fda-req...
[0] https://pmc.ncbi.nlm.nih.gov/articles/PMC1118461/
[1] https://www.faa.gov/ame_guide/media/AllergyAntihistamineImmu...
But that's not the important difference here. The important difference is that ceterizine has negligible antimuscarinic effects, unlike DPH, meclizine, cyclizine et al. Antimuscarinics are nasty drugs, and the antimuscarinic activity(and sometimes other non-histaminergic activity as well) is why a lot of first generation antihistamines are so bad for your brain.
Which moves us to the next two issues: liability and time. Any moment that you ask someone to revise a decision and specially with the stakes that the medical profession has that nobody has the time nor the inclination to open themselves for a mess.
Now, if you really want to be successful, you have to, before they even have a case with you, and specially before the diagnostic loop closes, to suggest the tests that the study has, since that has the biggest chances of looking at the right thing to look. Just be straight that you walked in with a theory. Doctors notice when they're being steered way faster than they notice when you're actually right. That's how you work with the systems that have a overworked mass trying their best.
My problem is that I needed information from 2 ENT visits to feed into ChatGPT to get that study. On the first visit he scoped my sinuses and immediately said "I can see evidence of allergic reaction, see those white bumps?". On the second visit I got an allergy stick test and it came out negative.
Those helped lead to that NIH study. It would have been very hard to have walked in with that study in hand.
> Let me just try taking a second-gen allergy tablet every day and see what happens.
Stupid question: Why did you wait three years before trying this tactic?Current Siemens MR software ‘Deep Resolve’ makes up the signal (adding about 50%), then makes up every second pixel, and then, for 3D sequences, makes up every second slice. It’s locking about 59% of the time off each sequences. And it’s really really good. I’m an MR tech.
After years of collecting artifacts and errors, I have more and more respect for the tool.
But it’s jarring. I open a sequence, decrease the acquired resolution, add the AI and get a scan that’s quicker and higher resolution.
It’s an amazing time to be an MR tech.
https://marketing.webassets.siemens-healthineers.com/2861d15...
No radiologist is buying "AI" scanners. Radiologists are probably among the most jaded of an audience about the word "AI" due to decades of undelivered promises. AI is synonymous with "worthless trash" to them, not to mention everyone says "AI" is going to put them out of work. lol
MRI is already a form of compressed sensing, I would much prefer statistical forms of super resolution to ones based on training data. Even if it is only trained on MRIs it will see some noise and plausibly expand it into whatever disease fits.
Actually, I'm curious what ChatGPT 5.5's ELO is- I wouldn't be too surprised if it's 2000+ just from its basic understanding of chess principles from all the content it has digested.
LLMs truly are marvels with text but anything spatial seems to really mess it up, somehow.
Not at all? LLMs are a terrible match for the kind of analysis a chess engine does (scaled deep search, deeply trained position evaluations). It's just not that kind of tool.
I've tried to pay chess with GPT-5.5, even played it again tonight, allowing it to use `python-chess` to keep track of the state of the position and to get a list of legal moves at each turn, so that it was fair. I also gave it blindfold odds, again to make it a fair fight, but it was not even close. GPT still isn't better than maybe 1000 Elo, maybe 1200 tops. Even with what amounts to being able to see the position and also being unable to make an illegal move, GPT-5.5 hangs material left and right, doesn't make a plan, and got smoked even when I gave it blindfold odds, to the point it's boring for me to play even under those conditions. I'm not sure it's better than whatever the GPT model was that was out about 8 months ago. I also thought it might be somewhat better than a beginner due to reading chess books, but no, it's complete garbage at playing chess, not even average-level skill.
Again, this is just one single person's experience. So not worth much.
I think we’ll see a lot of specialized VLMs that provide real value.
I don't understand why doctors don't prompt LLMs before saying wrong things. Is it ego?
I can understand for radiology because you need a specialized convolutional network, but for more knowledge based things...
I imagine reasons for what you’re asking might include:
* Prompting an LLM is work, and they’re already overworked just doctoring—every conversation with a computer is a conversation you’re not having with a patient;
* They’re probably right more often than they’re wrong;
* “When you hear hooves, think horses, not zebras”: the 15th case today of strep throat is probably strep throat, regardless of today’s 15th falsely-confident LLM weighing-up;
* They tend to have spent many many years honing a clinical intuition that makes an examination, to some degree, hard to articulate fully to the LLM;
* Liability/overdiagnosis: All this stuff is probabilistic. Inevitably, there’s going to be a time when the LLM throws out something I thought unlikely that turns out to be right, and there will be other times when it’s wrong but now I have to document why. How many false leads do I need to chase per one true differential? Does this really compare favorably to seeking a second opinion from another human doctor?
* Not everything needs to make it into the record. Once it’s in the LLM, it’s discoverable and litigable and hackable and permanent;
* Medicine is practiced in very different ways in different contexts—even in this thread, one radiologist routinely orders ultrasounds for soft tissue shoulder problems, and the other medical-world person replying has never heard of such a thing—presumably both within US health care contexts. Some doctors hand out antibiotics like candy, others are more cautious with respect to resistance. What’s right can depend on the time, the place, the clinical setting—more than just the immediate patient-level facts at hand, in ways that become awkward or unwise to express explicitly.
And of course… who’s to say they don’t do LLM-assisted research, in cases where they think it might be helpful?
Either that or laziness I'd imagine. This isn't limited to LLMs. Expert digital assistant systems that you query have existed for a long time. A good physician will double check anything even slightly unexpected against one.
you cant trust these toys at all. that doesn't make the useless, just untrustworthy.
Hurts who ? Yes the doctor is super stressed and has maybe 10 minutes for you that's the actual problem, it's not like before LLMs they were super glad to sit there and answer all your questions.
Studies have found that newer reasoning AIs are about as good at diagnosing illness from a written description of symptoms as doctors are.
Granted, it cannot actually examine a patient, so we're not replacing doctors anytime soon. But your view is obsolete.
It may have some utility after diagnosis, but this test doesn’t demonstrate utility for patients.
The more training data, the more questions it can answer with a reasonable degree of probability of accuracy.
Throwing away a potentially useful analysis just because it’s probabilistic seems a bit like throwing the baby out with the bath water.
This case is about handing a 3D imaging result to a text predictor and hoping for a valid second opinion.
The real question is where’s the cut-off point between accuracy and utility.
Remember: a second human opinion can also be wrong, and even a wrong opinion can still be useful (especially in medicine where differential diagnoses are a common practice - if the LLM gives you a useless opinion, you rule it out and move on).
I don’t think it’s particularly unreasonable to think that an LLM would have enough literature, or enough reasoning ability, to be able to generate a plausible interpretation of the data. A human can then review and say either “yeah that’s clearly not the case here” or “hmm, actually that could explain it, maybe we should order another test”.
But AI's problem is that its completely full of shit, sometimes, and the people most qualified to evaluate whether its full of shit are the doctors, not the patients, but just like OP's original article, patients are left feeling like their second opinion from AI might be more trustworthy than their doctors opinion.
Examples of things normal people can verify
- procedural errors that Claude can capture like some blatantly high dosage (grams instead of milligrams)
- outdated treatment plan, maybe there’s a credible new treatment plan that’s been used for years but the doctors were not updated
- literally being injected homeopathic drugs (takes no smart person to flag this)
Let’s stop talking as if doctors have a divine right here. And let’s accept some agency.
A doctor might have never recommended upping X, because they would know what it does to your body. Or they might have suggested additional supplementation to avoid this.
The fact that LLMs are trained on all public knowledge is a huge red flag, because there are more wrong infos out there than right ones. Especially about health, diet, etc.
It's now quite unusual that it's "Completely full of shit". If it contradicts something your doctor said I don't see why you should feel ashamed to bring it up. Sure it complicates the doctor's work, having ignorant obedient patients must be more comfortable for the doctor, but the end result could be more accurate diagnosis.
The same issues that were present with search-engine self diagnosis are still present with LLMs. If you provide Google with an incomplete list of symptoms and can’t interpret the information you find correctly, you will likely get an incorrect diagnosis. The same is true for LLM output.
There's a reason I ask AI about absolutely everything medical and there's a reason I keep extra quantities of prescription medications around for emergencies. I've saved my own ass a lot more times than the doctors have, thanks to good doctors not being available.
I get it. But the current system is also super difficult for the patient: getting time to ask questions, get clear answers, get the best possible diagnosis taking into account your history, symptoms etc and all that in 5-10 minute checkup when your doctor sees 50 patients a day and has very little time for you; this doesn't scale well. Patients run to A.I for a reason.
I wouldn't trust AI to make a diagnosis, but I would absolutely trust it to notice where procedure hasn't been correctly followed, where a treatment is counter-indicated because someone has missed a line on a health record, or where there's a clear potential alternate diagnosis which has been missed for spurious reasons. Also, unfortunately, where doctors aren't doing a decent job - often because they're overworked or underfunded.
And yea, I already did all the standard things. CBT for insomnia helped somewhat. My insurance didn’t fully cover it either, unless I was willing to wait for 8 to 12 months.
And I recently met someone with slow moving metastatic cancer. Thanks to LLMs they will most likely live another 3 to 5 years extra since the Dutch conventional mainline treatment hasn’t been taken yet. But it is German doctors that helped them and Belgian doctors that pointed out in a second opinion that a lot more can be done.
LLMs have a part to play. The false positives are awful, but I have seen an average of 5 out of 10 care when things become too complicated.
Except for trauma treatment. The Dutch healthcare system is amazing once they diagnose classic PTSD.
So it’s definitely not all bad but the trust I had when I was younger has been eroded quite a bit and LLMs can meaningfully step in, in my case at least.
[1] I know there are worse systems. But from what I have heard there are clearly better systems nowadays. It has slipped a lot
So 3 days out of 7 days I have guaranteed good sleep. The other 4 days are a toss up. But an average of 5 days of good sleep is much better than 3.5 days out of 7 days.
[1] https://www.kruidvat.nl/shiepz-melatonine-time-release-0-1mg... - Shiepz Melatonine Time Release 0,1mg Tabletten
I personally take 0.3 mg, two hours before bed. I've done this for about 2 years now. It still works. I know, anecdata, but as you can tell the dose is low.
Anecdotally, when I took mirtazapine for sleeping problems, it did sometimes seem to have a stronger effect the first time I took it after not using it for a while. After that the effect stayed stable. Overall it shouldn't cause habituation, and my doctor said as much.
Of course trust your doctor and not strangers on the internet, though.
Yea so this is where it gets murky for me. I experience some habituation actually. But my actual doctor went like "wtf is this?" and she didn't really mentioned what she knows about it. So on this particular pill my friend is my doctor. Not an ideal situation. I mean, he is an actual doctor but for him to be my doctor in this is a bit fucked up. He knows a lot more about mirtazapine than my GP though since he read up on it.
I suppose it can be quite different for different people. I stopped using it because it often (not always) made me still feel tired and unfocused in the morning, something that apparently also doesn't happen to everyone.
> But my actual doctor went like "wtf is this?"
Different country, but where I live, prescribing low-dose mirtazapine for insomnia appears to be fairly common practice even though it's off-label. I've had it suggested or mentioned by three or four different doctors, including GPs. I also know several other people who have been prescribed it.
The doctors here seem to prefer low-dose mirtazapine as safer over typical CNS depressants such as benzodiazepines for insomnia nowadays, at least if the problem may be longer-term.
So it's not really something particularly weird. Of course different countries also have different medical cultures so I guess it's not surprising if it's not that common in other places.
https://www.thecut.com/article/antihistamines-pepcid-ac-peri...
> Then, a few months ago, Angela saw a social-media post from a woman who took daily anti-histamines (like Allegra, Claritin, or Zyrtec) plus Pepcid AC (a common antacid) for her perimenopause symptoms. Her results, as reported, sounded miraculous: no more brain fog, no more tossing and turning all night. Even her mood vastly improved.
Instead of music, long podcasts you are given something to imagine at a time interval.
Like if you hear "calm river", imagine that. If you hear "heavy rain over a tree", imagine that.
In short → Close your eyes, listen & imagine.
[1] Account details when I wrote this down:
user: greybox555
created: 27 days ago
karma: 2The clanker said I'd be fine, I just needed some rest and OTC meds.
The medical staff immediately turfed me to surgery because the same set of symptoms I told the clanker were enough to concern them that I needed emergency surgery.
Had I have listened to the clanker, I'd be dead because I did need emergency surgery. (Hell, I almost kicked the bucket because I waited for someone to wake up to give me a lift because.my insurance probably doesnt cover an ambulance ride.)
Pretty much the like most manager these days, so I understand the frustration of the GPs.
We need studies that quantify error rates from each source type, then we need to account for the fact that the artificial type will keep improving.
The dad was a retired neuroscientist who delayed cancer treatment against medical advice because he was certain he had been misdiagnosed based on his own research that he did with the help of A.I.
https://www.nytimes.com/2026/04/13/well/ai-chatbots-cancer.h...
There's a comment on the article from Ben Riley:
> I am very grateful to Teddy Rosenbluth for sharing my father's story with the world, her kindness and curiousity proved to be restorative in ways I didn't anticipate.
> The two words that everyone used to describe my dad: "intelligent" and "kind," and he was indeed both of those things. The sad irony here is that it was his human intelligence, combined with these strange new tools that purport to be a form of 'artificial' intelligence, that led to his ill-advised decision to forego the treatment he needed for his CLL. A doctor has already commented on this story with the observation that AI "confidently asserts erroneous conclusions," and we simply have no idea how often this is happening or the magnitude of the harm that results.
> Not a day goes by that I don't feel the pang of my father's absence. He might still be here if not for AI. I try not to think about that, but sometimes I can't help myself.
Your comment is akin to saying "Karen from facebook who is a human pushed essential oils and ivermectin as a cure to cancer. Now doctor Y is suggesting chemo. Both are humans, humans cannot be trusted!"
This is the real root issue.
At 75 years old, he was stubborn. Is that reasonable ? Yes, perfectly. Could he have been right since the beginning ? Certainly. Did he deny evidence ? Yes.
Zero doubt that he was intelligent, everything points toward that direction, but that doesn't make a person less stubborn, because accepting the evidence, is also accepting that you were wrong if you initially postured yourself as adversarial instead of cooperative.
He would have read Wikipedia, scientific papers, etc, even without AI.
He did not want to be convinced. It works both ways:
https://www.foxnews.com/health/woman-says-chatgpt-saved-her-...
or
https://www.today.com/health/mom-chatgpt-diagnosis-pain-rcna...
Nonetheless, someone very smart, just didn't want to move from his position.
Like any domain, when you have questions or need a solution, you make research first, then you ask a specialist.
If you explain well the symptoms and context you can have proper advices and then decide on the path next:
Case A) It looks benign and advices / information that you collected seem reasonable, then you go your way.
Case B) You need second opinion of a specialist because the subject is too complex, or there are medications that you need approval.
Once you have challenged LLMs, and read about the topics over and over then you genuinely become really good at understanding it (especially if you triangulate over LLMs and ask them to challenge, you start to have genuine questions). No matter if the answer is right or wrong, you have elements. Maybe you missed the point, but you come prepared.At home you have the time to assess the options, pros and cons of each approaches, the possible questions to ask and then challenge the doctor.
Shared decision-making is an actual evidence-based model of care, and patients who arrive understanding their condition and carrying specific questions tend to get better attention and better outcomes.
Some doctors get annoyed, because they have big ego and choose to be patronizing, but it is exactly their job to answer such questions.
With LLMs, it's quite good, you get nuanced and rather useful answers.
Before LLMs, no matter the topic you searched for, the answer was the same: "you have cancer / an [obviously deadly] rare disease"
The other problem, in many places: • The doctors are not affordable
• They are too busy for you (< 15 minutes)
• You may need to wait months to get an appointment
• They are not good (country-side is an example, and sometimes even country-level)
+ you can have all of these factors together.So, you have something deeply bothering you, your only appointment is in 4 months. It would be insane not to take the time to explore different solutions and not to come informed about the topic.
If you express your prompt properly and do not rely on imagery, you can absolutely have top-tier advices.
I told my mechanic the film flam is broken but he said it was the rim ram. He fixed it and we all went in with our lives.
But doctors insist on this God like status so it’s a “nightmare” when patients try to help themselves.
A con artist, a fraud
[0]: IF.
It's a 180 for me: While I believe doctors should explain diagnosis or treatment decisions when asked, I don't believe they should be taxed with explaining away alternatives. In my anecdotal 2nd- and 3rd-hand experience, doing that is taking at least a third of their time (on roughly 5% of the patients who think demanding answers will make things better) -- with zero improvement to diagnostic accuracy or treatment effectiveness. Doctors already consult with other doctors, and it makes no sense for them to have to consult with ignorant patients or treat their AI psychosis on top of their disease. It doesn't increase patient autonomy any more than adding a steering wheel for child car seats would help toddlers learn to drive.
Dr. GPT is a good brainstorming tool. It helps synthesize information in a way that primary texts don’t. But it does force you to say “that doesn’t make sense”.
I do think that people saying “doctors don’t know the state of the art” have a weaker case. If you think about it in terms of token density during pretraining and how post training datasets are constructed, I think it would take us a very long time to adapt to any fundamental shifts. If we have forgotten how to cure scurvy, how many journal articles would it take before we adapt to a discovery?
This is kinda the case though. In Poland I met only one psychiatrist that knew about DSM-5. In this year. DSM-5 was a thing from 2013.
Doctors are people just as us, not every single of them is good.
It is kinda spooky, though, to have freshly minted doctors from a few years back whose school-knowledge will forever be "outdated and archaic" based on standards published before they were in school.
Some good advice I got: treat this as a generation shift, find younger and newer doctors who are familiar with the "modern" standards.
Churn is built in to the specialty. Read into that what you will.
[1]: https://www.apa.org/practice/guidelines/criteria?item=5
I would not trust my health to LLMs, that is true.
Incidental Rotator Cuff Abnormalities on Magnetic Resonance Imaging https://jamanetwork.com/journals/jamainternalmedicine/fullar...
Claude: Primary finding: Complete ACL tear with the classic pivot-shift bone bruise signature (posterior lateral femoral condyle + anterior lateral tibial plateau edema) and large hemarthrosis. PCL, MCL, LCL, menisci, and cartilage all intact.
Radiologist: English translation of findings & conclusion: Mild joint effusion. No Baker's cyst. Post-ACL reconstruction with minor cyst formation in both the femoral and tibial bone tunnels. The ACL graft shows heterogeneous signal but no complete or recurrent rupture. PCL and collateral ligaments intact. The lateral meniscus appears abnormal, likely from prior partial meniscectomy, with significant cartilage loss (partly Grade 4) at the posterior lateral compartment, osteophyte formation, and reactive bone marrow edema. The medial meniscus shows diffuse signal change from prior repair but no recurrent tear (specifically no recurrent bucket-handle tear). Mild chondropathy with focal cartilage loss on the lateral side of the medial femoral condyle. Cyclops lesion present. No definite loose bodies.
It would already be a huge benefit to 90% of people worldwide if the very first part of most hospital visits would be outsourced to frontier-level LLMs. Yet this kind of misuse just gives the medical industry a stick to beat that idea into the ground.
Oh well, I'm sure there will be at least a few countries that will indeed embrace frontier models for initial diagnostic medical purposes. Maybe medical tourism destinations. But it's unfortunate for those who can't afford the trip.
There are other commenters saying this is a good practice they've also done for other injuries. You are saying you are an actual radiologist and immediately clock the problems with its advice.
I have seen this pattern over and over again. Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading. It is only when you do not know what the AI is being asked to do is it likely you will find the output helpful.
This is itself alarming to me, but no one else seems to find this to be quite damning for the AI services being offered, preferring instanced to be wowed by the convenience and speed at which they can be delivered unreviewed and unproven information.
Yes, this is exactly so. AI is able to confidently sound plausible enough to convince laypersons or anyone who isn't very familiar with the subject matter, which is a big part of the mass-appeal "magic" of ChatGPT and other similar tools. It's like having a know-it-all friend (who also makes shit up to bridge their own knowledge gaps).
In many non-advanced non-specialized situations, AI is right enough to be at best useful or at worst not harmful (usually landing in the middle somewhere).
But speaking for myself, in areas where I consider myself quite proficient, I can very easily spot the subtle inconsistencies and naive conclusions that AI responses provide, and I have to guide/steer/correct it a lot to get good results when the subject matter is complex enough.
"Be wowed by the convenience and speed", or merely "take advantage of the mere availability"? What most people find to be damning about expert advice is that they simply can't get it anywhere, at any cost that they can afford.
Who do you choose to be coached by an expert on the ground?
The first: Has no clue about anything and therefore no useful knowledge and cannot challenge me
The second one: Is proven to willfully give wrong information and will make me do mistakes for sure.
The LLMs will do their best, even if imperfect, since they summarizes what appeared in books.
I prefer to be grounded on what Airbus / Boeing manuals, or on what pilots training book said, than two far more unreliable sources.
Properly emotionally processing this fact and your complete inability to do anything about it is called an "existential crisis" and if you haven't had one or several yet, you will.
Putting that aside, your philosophy sounds shallow. Death is certain, but how long you have to live and the quality of that life are not predefined. An incompetent passenger-pilot trying to save you from a crash will at worst make no difference. But an incompetent doctor can teach you that death isn’t necessarily the worst outcome.
I think the different ways people accept death explains a lot of people's psychology, like how you can guess people's attachment styles or Freudian stage fixation. For instance, billionaires who pour all their money into anti-aging research clearly are not handling it well.
Ok for pain in your shoulder it might not, but how about a woman with a lump in her breast waiting for the mammogram interpretation? How about someone trying to understand disturbing lab results? People are also often pushed these days to move through visits with doctors at a breakneck speed, but the AI will "hear you out" all day.
Part of this is a problem with the AI, part of it a problem with our healthcare systems, and part of it is simply human nature. If you think that OpenAI, Anthropic, Google and the rest weren't aware of this going in you must have very little faith in the intelligence of their members. It's not hard to imagine the future of LLM's should involve a hell of a lot of liability on the companies running it, but for now it's the Wild West.
Whatever scenario you come up with my answer is the same.
As an adult I’d like to be able to choose what tools I use to learn about my condition regardless of how well it works or even if it’s likely to mislead me.
There’s risk in every aspect of life and we can’t baby proof everything.
Even if it "works" so poorly that you're not actually learning about your condition?
So if you MUST have answers that are at most random guesses, I'd suggest saving a few bucks and asking a coin before flipping it.
Current trend is that the models will try to explicitly steer you towards "asking better questions from your medical provider", rather than providing diagnoses. They do also evaluate whether something can actually be established rather than just listen and nod along. And so the "you must have very little faith in the intelligence of their members" goes right back against these failure mode ideas.
Now of course, given a sufficiently desperate person, they can probably torture anything they want to hear out of these models. But so can they out of actual people, so that's kind of a high bar. When you get to the point where people are willfully misreading a given piece of text, bets tend to be rather off.
AI isn't even the first instance of this phenomenon, news articles are like this as well.
It is weirdly religious in a way, because if you were to present contrary evidence (e.g. experts in a field weighing in about how plausible sounding responses are bunk), you would only be told you don’t believe enough in the long term potential and capabilities.
Don’t get me wrong, I think we all agree capabilities will eventually improve (and farther-future capabilities could reasonably surpass experts), but really is unclear if the current transformer architectures with their probabilistic/hallucinatory outputs will plateau before they surpass current experts abilities in all promised fields.
A lot of the models up to this point have been benefitted - like Google did - from essentially ‘pre SEO’ internet.
Now the same tools are being used to generate nigh infinite good sounding bullshit, which poisons the dataset in all sorts of hard to detect ways.
To add insult to injury, the human experts are also not as. Naive, and have many incentives to poison their own input in subtle ways too.
OpenEvidence claims
"More than 40% of U.S. physicians use it daily, and it handled around 20 million clinical consultations per month. Over 100 million Americans were treated by a doctor using it in 2025."
https://www.cnbc.com/2026/01/21/openevidence-chatgpt-for-doc...Here is an example. My provider sent me this note. I'm quoting verbatim here from my MyChart record:
"Your liver enzymes are high, I would like to order acetaminophen containing medication like Tylenol, I would like to order liver ultrasound I placed ultrasound order in the system, make an appointment for radiology, I would like you to get hepatitis panel lab work done, obtain blood work order, please schedule a well visit to get it done"
When I queried it, this is what I got back. It was a dictation error. You could almost hear the panic in the message:
"Sorry for wrong message earlier, I was dictated message- so could not realize that it was written to take Tylenol type of medicines- I DO NOT RECOMMEND ACETAMINOPHEN CONTAINING MEDICINE - LIKE TYLENOL AND ALCOHOL DUE TO ELEVATED LIVER ENZYMES."
Again the problem is not dictation, or LLMs. The problem is humans ignoring their responsibility to check the output of a machine.
100%. Also, management.
I wish someone would go ahead and coin an AI version of Amdahl's law that states the work speedup from AI is dependent on amount of unverified AI output used.
Iow, if you 1:1 verified everything, there would be no time savings.
Ergo, you get management saying (1) we demand time savings due to AI & (2) we demand you fully check anything you use AI for.
End result? People skip (2) to hit (1).
Then management burns anyone at the stake whenever inevitable mistakes happen.
Can you think of any real life examples where an LLM is likely to be used?
I think in practice what you're saying is there are problems where there exist efficient deterministic verification methods, and I'm sure that's true.
But that's not the bulk of everyday work LLMs are being asked to do nowadays across industry.
But if you want to keep it in the realm of the everyday: you're asking if it is easier to write an email than to read it and check it covers what you wanted to say? Is it easier to search for something or to look at what's been found and say that it's what you were looking for?
Which means she ends up spending just as much time as if she’d done it herself as it needs to be verified for accuracy every time…
If a physician uses Google to search for a dosage chart for some drug they rarely prescribe, you wouldn’t say they are using Google to diagnose the patient. You wouldn’t say that either if they used Google to search for the most recent studies on a topic.
The fact that they use it doesn't make what the result is any worse or less trustworthy - arguably it makes it better.
It only becomes a problem if they offload all of the thinking to AI.
For one, if your website/book is poisoned, who is going to trust it for anything at all, much less for training models?
For two, all the major AI labs hire or contract for subject matter experts to create curated data sets, evaluate model performance, etc.
Unless they hire malicious experts, this will provide a growing, high quality data set that should drown out any poisoned pretraining data.
Yes AI scrapers can easily spoof user-agent, but they fall out of date as the browser updates.
Bit harder to catch them in tarpits and then serve nonsense to whoever ever triggered the tarpit.
It’s a hell of a lot easier for a company to ensure that its scrapers all report the latest user agent string than it is to get everyone and their mother to update their browsers in a timely fashion.
and browsers forcibly auto-update
This is how we get LLM summaries presenting something mentioned once by some nutjob in a reddit thread as bona fide FACT
If it's easy enough that some randos can do it for fun, what do you think happens when there's commercial interest behind it?
Obviously companies are going try nudging AI towards recommending whatever they're selling. It's a logical extension of SEO - and that's a 100 billion USD industry.
Additionally, if I believed myself to be in some sort of spending - err - AI race, I'd try to poison the data sets of my competitors by putting crap out there for others to ingest.
What does it mean, Is it like when somebody used some coding agent to develop a feature and later input prompts and a resulting PR can be used for training by a presumption that final PR was a correct implementation of a prompt?
The trick is to find the examples that are just in between too difficult and too easy for the existing agent, these have the strongest training signals
Peer reviewed journals, textbooks, in-house teams of experts, trusted news publications, etc.
The whole idea of scraping large swaths of the internet for training data has always been pretty dubious due to the variable data quality.
I mean, just look at the early Google models that told people to put glue in their pizza due to a joke in the training set. Garbage in, garbage out.
This is one of the first and most obvious problems all of these labs have run into, and countermeasures are only going to improve.
Also, those other sources are getting buried in AI slop too.
Your original claim is that this will be enough of a problem to prevent models from improving in expert level knowledge. I completely disagree with this premise.
If the models fail to improve, it will likely be due to limitations in the transformer architecture rather than poisoned training data.
And even then, I doubt that the transformer is the best architecture we will ever come up with.
Clearly it doesn’t learn or think like a human does, since humans don’t need many gigabytes of text samples to learn to talk, so there is some room for improvement.
What I’m saying is that the AI labs are handling this not by fixing the “garbage out” part, but by minimizing the “garbage in” part.
The fact that all you could come up with was research (not an actual example of poisoning a real training set) from 2025 kind of proves that this isn’t some kind of widespread, unsolvable problem like you seem to be claiming.
The poisoning issue makes it so that no one can use the internet for training anymore, because more and more internet content is poisoned as a side effect - or poisoned intentionally. And .001% of poisoned data is enough to screw things up if included in the training data.
It’s also one reason why Google search results have been getting so much worse - it’s hard to not find a SEO page with subtly (or not so subtly) wrong AI slop on almost every topic you can imagine. Most folks won’t recognize it, but that’s what is going on if you know what to look for.
One other way of putting it is the ouroborus problem - more and more internet content is AI generated, because of people trying to game the system, and they are making it is indistinguishable from real content as possible to get by the AI detection algorithms.
Anyone trying to train on it just ends up eating the shit from another LLM, which poisons it.
Another name for it is ‘model collapse’, which also doesn’t have a known solution yet.
And do you even realize how much data 0.001% of the training data for a frontier models is? They’re trained on 10s of trillions of tokens, meaning you’d need hundreds of millions of tokens of poisoned data.
Some of these problems you mention could become real barriers to models improvements, though there are plenty of countermeasures, such as by focusing on high quality data sources like I mentioned before.
We’ve already probably gotten as much as we’re ever going to get from simply scraping more and more unstructured text from the web as a way to improve model performance.
The type of training being done now is around tool use and solving specific types of problems better, which is the type of training data you simply don’t find lying around on the web.
This is exhausting.
It’s like arguing crypto with someone who has never actually committed a line of code. Why do I even bother?
I’m wondering the same thing. You keep talking of some grand poisoning problem but can’t point to any specific public information except an article saying that it’s possible. As if that was ever in doubt.
Guess we’ll just have to agree to disagree.
An expert already knows they don't know everything. That was never the point. Critical thinking cannot be delegated to AI any more than it can be delegated to a book. There is nothing new going on here.
And it's so much like listening to someone in a church congregation sharing their experiences with god. Clear and obvious gaps are hand-waved away exactly how you're describing.
While I can understand being skeptical of non-experts' claims that such answers are enough, I don't understand why you call it "psychosis" and not simply naivety or lack of expertise.
At the same time, the new so-called "models" haven't been pure transformer-based LLMs, but entire systems with tools (with access to the Internet), data storage, and the options to trigger additional instances for different tasks.
"Oh you like LLMs? You must in AI psychosis!"
Let's not pretend it is anything more than the run of the mill wet fart of a culture war label. It's quite literally the "TDS" of the anti-AI crowd.
The idea here is to signal that you can absolutely use LLMs to help you figure something out. But also, they're wrong a lot. So use your own brain too.
The problem is that AI psychosis is fundamentally the belief that an LLM is "thinking" at all. Outputs are just believable word vomit which resembles factual information.
The problem is real but I don't think positing a philosophical root is helpful
If "agency" is making decisions and performing corresponding actions in the real world, then LLMs most definitely LOOK LIKE they're making decisions (what's the next token? which tool to use? what's to say, in general? what idea to convey?) and performing actions (tool use). Can we tell whether they are ACTUALLY making decisions? Well, are the people around me "actually" making decisions? Or are they simply pushed around by circumstances and external forces?
Am I actually making decisions? Did I like DECIDE to write this comment? Maybe? I have no clue...
It's quite simple, the agency that the LLM appears to have is actually your own. Without a prompt an LLM does nothing. It has no thoughts between prompts about you or your problems.
Twice in your comment you suggest things that you think that I believe, please do not do this.
So when it's not active, not responding to a prompt, it's of course not thinking. I'm pretty sure nobody actually questions this. Is your computer "thinking" when it's powered off? Can a piece of metal think? Probably not. So there are no thoughts between prompts, this seems obvious.
Thus, this is a question of "discrete time vs continuous time". LLMs "live" from prompt to prompt. Humans are alive continuously. In some sense, we're prompted by a lot of things all the time. As I'm writing this, I'm seeing stuff, I'm hearing stuff, I can feel various parts of my body, I'm thinking about my problems, my goals, other people's problems and goals, etc. When I'm in a sensory deprivation tank, my brain keeps "entertaining" me by "self-prompting", like a recurrent neural network (I guess it literally is a massive RNN).
So it seems like your definition of "thinking" hinges upon the LLMs being discrete-time and single-threaded (can't think about multiple things in parallel).
IMO a more interesting question is whether an LLM is thinking WHILE IT'S GENERATING A RESPONSE, while it's "alive".
> Consciousness isn’t really guessing the next thing to say-
I don't know what consciousness is either and these debates are a dumpster fire when they happen, but it sounds like you're pulling forward this "LLMs are just predicting the next token" (true by construction) implies that they can't learn or reason or be conscious (2/3 are wrong, the last one isn't falsifiable without a useful definition).
And context window work very well. You can 'teach' an llm a new programming lanuage and other things through it.
You are anyway, I don't see anyone up the chain saying that.
Do you think it is any more possible to have a proper discussion with someone who preemptively paints the other person as mentally ill? Or someone who preemptively victimizes themselves?
Cause I don't think these are the hallmarks of an honest discussion. See also the entire past decade of political discourse.
Like, consider this:
> It is weirdly religious in a way, because if you were to present contrary evidence (e.g. experts in a field weighing in about how plausible sounding responses are bunk), you would only be told you don’t believe enough in the long term potential and capabilities.
A trivial counter to this is that you can just be an expert at something (e.g. your own work), use the damn thing yourself (professionally), and evaluate the outcomes for yourself. Then maybe remark "LLM good".
Now you come and remark "LLM bad", and point at random "evidence", either of outright other workloads, or even the one at hand: you're asking someone to reject the reality they've already experienced, entirely based on the assumption that they're "merely religious" or "in psychosis". You tell me if that's any more epistemically rigorous and sensible than their story.
But on your actual point, I don't think AI needs to "surpass current experts abilities in all promised fields" as a marker of its ability. The immediate gains has already shown some remarkable promise and more LLMs should have had safeguards around mental health up front. If I were to put it on a scale, I would say it is net positive long term with a strong negative spike up front which was somewhat preventable. But who knows, maybe is just have "AI Psychosis" and you can easily dismiss me.
Similarly with LLMs, you can't just write them off entirely because they sometimes provide misleading or incorrect advice. The positive utility maximizing view is to learn when you need to call in an expert. I recently moved in to a new house and have used Claude extensively to figure out basic things (e.g., adjusting the garage door height, how to mount a TV). However, when the HVAC suddenly stopped working, I gave Claude a shot for an hour and tried some non-destructive fixes, but then realized I had to call in an HVAC expert.
I find Claude is surprisingly similar to a confident but incorrect coworker, with the benefit that Claude will reevaluate when I correct it.
I guess to me it has to be comparable to be an alternative.
Like, I don’t consider doomscrolling x an alternative to reading Wikipedia but I might consider it an alternative to CNN, even though they’re all technically and very broadly activities that I could use to inform myself.
In that same way I don’t consider the multitude of ways I could use my free will necessarily alternatives to each other even though they technically are. It kinda sucks but going that broad feels to me like it breaks the concept of alternative and makes it kind of meaningless.
Then we had computerized encyclopedias and search engines that searched the library.
I mean, you had to work for the knowledge. Sometimes you didn’t know something and no one else knew either, so you had to wait until you got a chance to find out, but you would think about it and sometimes you would be right when you found a reference source.
I’ll also note, Wikipedia is a secondary source. It is not a reliable source of truth. It is more like the ‘ask someone else’ alternative than anything else, it’s just ‘someone else’ is a person on the internet who writes Wikipedia articles.
I'm seeing this fairly often and when it isn't garbage it's a capable person who has gotten inspired by their 'collaboration' in which the busywork is being done by a machine, but they're doing so much directing and correcting that it's not unlike what would happen if they got heavy into meth and went on a tear.
You absolutely can write them off entirely and decide for yourself what your comfort level of human-killing speed-freakism you want to pursue in your productivity. There's a long history of humans managing astonishing levels of productivity through self-destructive means. This is not even cheaper, once the 'first one's free' wears off: it's just a novel method of getting humans to burn themselves harder in the belief that they have a magic feather.
The ones who're really throwing themselves into the situation are the ones who'll burn out, but who aren't setting themselves up for atrophy and learned helplessness. Anyone who believes the technology lets them be a lazy manager just getting paid, is in for an unpleasant discovery.
For example, we had to advocate for certain practices during the birth of our first child that became routine during our second several years later.
So, neither side is guaranteed correct, doctor or citizen researcher (which did not include LLMs in my case, for the record). The truest answer is also the most useless one, applicable to all fields: it depends.
The real question is: if you embrace being a layman, whom do you trust more: LLMs/the internet or experts, like doctors? I think the answer is pretty clearly experts.
I.e. nothing this radiologist said was related to the LLM’s advice.
Then to say "Aha, but all of that is AI psychosis" makes obviously no sense: Why would we trust experts when they offer critique but not when they say "this is helpful"?
Overall: People are not insane. AI makes mistakes and, often, fails completely. AI also helps them do things better, quicker, increasingly so. The jaggedness of AI is confusing and real.
There is a huge difference between having a chance of a good result, which can be useful for experts able to filter out the bullshit, and consistent success. I would generate code as a helper, I would never allow a guy from marketing to merge unreviewed AI code.
As an industry we've been promising people for decades that if they put all their data into our special softwares they can get all sorts of information back out that will make life easier for them, reveal new insights and otherwise improve their understanding. But the unspoken caveat has always been that you have to put the right data into the right places, in the right format, in the right way and then you have to ask the right questions, in the right syntax, with the right tools. And if you get any one of those parts wrong, you're not going to get the right answers (or possibly even any answer at all). How many people have had their excel worksheet that they (or someone else they asked/employed) built for some task that has been working fine for the last year suddenly stop working or start throwing out nonsense numbers because some input changed? Or how many people have experienced their system seemingly throw out meaningless garbage because daylight savings changed right at the moment the report was being run? Or spent months operating on wrong data because the person who wrote the query misplaced a parenthesis and the query was searching for "(foo AND bar) OR baz" and not "foo AND (bar OR baz)". For most people, the computer and the programs they use to do their jobs are magical black boxes that most of the time produce mostly the right answers and sometimes get things very very wrong with no indication of what has changed. Which is effectively the same experience they will have with an AI, but now instead of needing to figure out some arcane excel pivot table and VBA script, they can just dump some raw data and a "natural language" question into the AI.
And that's not counting the fact that their experience with looking information up online is about the same as well. How many absolutely confident wrong takes have you encountered online for things you're an expert in? How many of those wrong takes have come straight from supposedly trustworthy sources like news companies or even other people in the field?
For most people, using a computer has always come with the asterisk that you should always be aware that the source you're reading could be very wrong, that the output is only correct assuming all the inputs and all the parts processing that input are also correct and that everything you do should be accompanied by vetting by experts, whether those experts were software developers or domain experts. For most people the only thing that's changed with AI is that it's a one stop shop for their "probably directionally right, almost certainly wrong in the details" access to the digital oracles.
But see now we are talking about something else entirely than the claim that I found dubious, which was: "Anytime someone is an actual expert at anything, AI output appears insufficient or incomplete or outright misleading."
Consistently good enough !== anytime insufficient
Apply that to the Internet at large, and realize where LLMs got their training. They're basically ConfidentlyIncorrect personified.
The LLM may have, from its "perspective", implicitly thought the OP was telling it that he had strong reason to believe there was no calcification and was not considering the bigger picture of possibly receiving an incomplete/poor assessment from the medical staff. In fact, the issue here may be the LLM overly trusting doctors vs. trusting its own expertise.
We've known since the beginning that AIs confidently say incorrect things. But now that they can speak confidently about very complex topics, and mostly say correct things, we are letting our guard down and lots of subtle falsehoods are slipping through.
*In one case, I was able to put things back on track because the AI suggested my colleague talk to me; somehow it figured out we were co-workers.
Absolutely agree. Have seen this first hand
Welcome to the club? This new awareness you've found over the true quality of LLM based GenAI output has been what "all the haters" have been mad about for-ever. That the output of LLMs are clearly defective, and merely have found a cute trick towards making humans think they're less defective than they are actually measured to be.
And the corresponding anger and frustration to push the risks of genai output out onto others, while also aggressively pushing it as a feature you should be using already. You're behind don't you know, and whatever other lie I have to tell to trick you into enough FOMO to pay me 200USD/mo so I can sell FOSS back to you.
An LLM can only output the mean next likely token, and then add a bunch of extra noise on top of that so it feels interesting and not repetitive. None of this is new, the problem is, 50% of humans are below the mean, but have no idea. So when an LLM tells them some lie: well, it sounds so helpful! It's impossible for someone who sounds this helpful to lie to me, liars never sound confident! It must be PERFECT! I'm gonna tell everyone how perfect it is. so the bottom 0-33% think LLMs are fantastic tools that make nearly 0 mistakes in comparison to the bottom 33%. 33-66%-ish aren't sure, some times it's great, but it will make that random mistake sometimes, but I can catch most (or all of them depending on ego). and the 66%+ are angry about how many people are getting tricked by something so obviously low quality, or are lucky enough to not have to care.
So when an LLM was asked to analyze the unit distance conjecture, it just spat out a bunch of average-or-random tokens that coincidentally happened to correspond to a valid proof that had eluded humans for decades?
yes
https://en.wikipedia.org/wiki/Texas_sharpshooter_fallacy
How many problems and/or times did it make up completely random bullshit with no basis in reality? Random noise looks really cool or impressive when it's right, but if and only if, you're willing to ignore all the times it was wrong.
I'm not.
More on topic: if the article's author arrived at a definitively negative result would this have shown up on HN?
This is completely different than asking for general medical reasoning which is more derived from papers, public standards and textbooks.
Text exists at the right scale but images don’t.
The term for when the press "gets it wrong" is Gell-Mann Amnesia (https://en.wiktionary.org/wiki/Gell-Mann_Amnesia_effect).
In that case, when you have personal knowledge of the facts, or know the specific domain area, you can see where the reporter mixed things up.
AI is no different, it's just a bunch of matrix math substituting for "the reporter" regurgitating what it was previously told. So the Gell-Mann Amnesia effect would apply just the same. If you have domain knowledge, you immediately see where the AI got it wrong. When you do not have domain knowledge, you have less chance of seeing where the AI was wrong.
I always recommend people try asking LLMs a lot of questions on something they know first. Programmers should start by asking LLMs to work on a codebase they’re familiar with first.
You’re overstating the problem, though. Even for an expert the LLM will get a lot of things right and can be helpful under a watchful eye.
The real problem is knowing how to identify when it’s on the right track and when you need to correct it, because both cases are presented with the same tone and confidence.
An expert can better identify when the LLM output doesn’t sound plausible. Someone unfamiliar with the topic will think everything it says looks correct.
It has been like this since the rise of "AI". The only people enthusiastic about it are usually the ones hoping to make a profit in one way or another.
In fields where I'm an expert... it makes a lot of silly mistakes that are annoying and I feel like they would just cascade if I didn't correct them early. (I still think it's a net win, but... I watch it and it watches me, and we both do better work. I'd even apply the "magical" adjective when it does stuff I hate but know how to do, like edit Helm charts. What would normally be 20 minutes of me griping about YAML indentation is just a correct diff in seconds. I'll take it!)
So with that in mind, I tend to distrust output that I can't verify. If a doctor was recommending surgery and I thought the plan was too aggressive, I'd get a second opinion. I don't expect Claude Code to have much medical diagnostic ability, as that is really not what the model is trained for, and I know how it performs on work that it's trained and fine-tuned for. That is not to say the output is wrong and that it can't have diagnostic value, just that I personally wouldn't feel safe trusting it. Wrap up the same model with fine-tuning in the domain and a harness that reminds Claude to do a lot of sanity checks, perhaps with a human in the loop to guide it back onto the rails when it gets hyperfixated on something that doesn't matter? That could very much be a useful AI product.
1. operates against a well-defined DB of medical studies
2. intakes my basic demographics, vital signs and medical history
3. quantify uncertainty wrt a specific diagnostic (it's own or one received by a healthcare professional)
4. specify medical tests that can be executed, and how they can be obtained
5. provide scripts for interacting with healthcare professionals/functionaries
I would imagine that the thing would need distinct operating modes:
1. A diagnosis generator
2. A diagnosis evaluator/critiquer
3. A patient educator
Software is one domain where it excels because of structured training data and simulation environments, so I'm well aware it's better here than other areas.
Still there's somewhere balanced between saying every time it's "insufficient or incomplete or outright misleading" and "just trust AI". AI's a useful source of information/reasoning/research, but know you need to validate it's answers for important decisions.
AI is much worse.
AI assistant are industrializing the Gell-Mann amnesia effect.
media is awash at the moment with experts chiming in to support AI, saying their fields are being revolutionized, etc.
it seems unsurprising to me that the laymen opinion would follow the loudest media trumpets.
A real doctor is accountable.
They might both "know" a lot of things but implicitly the party who is accountable is going to be more trustworthy.
And I don't see that going away until AI companies must be licensed for application x and can lose their license / be sued if engaging in malpractice.
Do these LLMs make mistakes? They sure do, I see it all the time. But they can also help people make breakthroughs.
And this isn't the only time that Gemini has helped me diagnose long-term health issues, either.
I am not advocating to trust anything they say blindly, but they can be a great place to form new hypotheses and learn the right terms to look for when you are unfamiliar with a subject.
After a couple of rounds of that, a picture will start to emerge. The AI will make a few XYZ hypotheses of what may be going on, some of which will make more sense to you than others. This is when you can start searching some of those terms in places like pubmed.ncbi.nlm.nih.gov, including for example like diagnostic criteria for XYZ.
One of the ways I often use these AIs, not just in the context of finding possible diagnoses, is requesting them to make the case for and against hypothesis XYZ based on the data you have personally collected. Again, it's not about fully buying every thing that comes out of them, but it can help you consider angles or possibilities that did not occur to you, or that you had previously accepted/discarded without sufficient evidence. Think of them as that quirky acquaintance that knows a little bit about everything but sometimes misremembers, rather than as a god-like oracle.
And don't do all this in a single session/context. Start a new context every now and then, because otherwise it tends to go in circles as these AIs are biased towards agreeing with whatever it is you said most recently. Intentionally challenge yourself, re-evaluate the existing data from other perspectives.
Sometimes what you learn is not pleasant, but as more data becomes available, you learn to accept it. Good luck.
I have seen outputs that look good but the actual content is bad. If you’re inexperienced in a field you can’t see it because AI makes anything look right.
I have gotten very good results with AI but you can’t take the first answer at face value. You need to be suspicious and challenging until you tweak out the right answer over time.
https://www.nature.com/articles/d41586-026-01947-1
I've started asking my doctors whether they use AI, and if they say yes look for another one.
A very plausible explanation for the adenoma detection rate to have gone down is simply that its prevalence went down among the population in the second three-month period.
This was not a randomized trial. Concluding that "AI usage degrades physicians' skills" is questionable at the very least.
https://www.sciencedirect.com/science/article/pii/S245195882... (+ cf. its references)
One doctor diagnosis + LLM is gonna throw you off. You need more datapoints.
I wonder if this person was going to a traditional doctor or if they were visiting some type of specialty clinic as a second opinion. For most conditions you can find specialty clinics that will prescribe and administer (and bill for) a lot of non-indicated treatments, but some patients like being in the care of doctors who take action and do things after being recommended more conservative treatments by primary doctors.
1. stop any inflammation, by taking NSAIDs for a few days
2. detect and correct any behavioral patterns that could have caused the presumed overwear of the tendon
2. start physiotherapy to strengthen those muscles that can take over the load from the damaged tendon
These are not quick fixes, because quick fixes don't exist here. Stuff like shockwave treatment, massages etc will only lessen the problems for a few hours at most, after which they will come back.
Well, we now have the best model of our time (trillions of $$$ of investments) telling us something completely different(and wrong) from a human expert. I would really like someone calling out dario, sam, elon on these things and hear their explanations but alas, a man can only dream.
diffusion models are probably a better bet for identifying irregular structures
I think they’re artificially stunting the field to raise their wages. For example in my city the medical school only accepts 11 people into the program a year. (With an average graduation rate or 3-5). My niece has been trying for 2 years and finally got in this last year. Even radiology is doing AI assisted diagnostics. Half my MRI’s from this year has Doctor notes and HealthBot (AI) notes attached to them.
~ I’m assuming other schools severely limit their radiology admissions as well. To keep the wages high and the field desirable.
These days Xray machines - they don't even suit up in lead or stand behind a wall , just point and shoot. In fact they're nice and portable. I wish i had a xray machine at home.
Funny how the jobs most at risk of automation now are tech jobs.
So, unless you can turn the image into a natively tokenized format like JSON or something that somehow accurately tokenizes what's on there, I would NOT trust Dr. Claude's analysis. If you want a second opinion, talk to another doctor. A human doctor.
> AI can absolutely shatter that feeling in an uncomfortable way ...
I see this as a field report in a time of fundamental transition, from a world without AI, to one that accommodates/incorporates AI. For this to happen, AI will need to become more trustworthy. As for the U.S. medical system, it can't get much worse.
I recently had a similar experience (meaning walking a fence between old and new methods), where I was told I could get an appointment with a human medical practitioner in nine months. So, to resolve my anxiety I consulted AI and got an instant diagnosis, one that was later confirmed by the inaccessible medics.
Being a born skeptic I wasn't going to act on AI's diagnosis, I just wanted to know what was going on, resolve some uncertainty. Another advantage: an AI chatbot doesn't say, "Wait, you're on Medicare? Hmm. See you in nine months."
Don't take this as an endorsement of AI's diagnostic abilities -- it's way too soon for that. In my case it was a slam dunk, about a condition I knew nothing about.
I didnt see the full process but I used unet models for tumor detection so I am somewhat familiar with the possible caveats of any evaluation from a engineer perspective.
First, I would like to point that unfortunately, it is not uncommon to go to two different human doctors and also get two unreliable diagnosis and treatment. The biggest problem, in the way people plan to use ai on health is the lack of liability.
A bug on a regular old web site doesn't kill anyway nor cause pain and suffering (most of the times) but misdiagnosis + the fact that a model is very good on presenting arguments even when it is completely wrong.
Claude code, and I am talking about opus 4.8 here, can tell rivers of information about code pattern and develop the poopiest code the next line.
This is a machine that will deliver a sort of templates document based on the input information but it is not exactly doing the work if you don't directly it to do it right constantly.
Because the model isn't thinking I wonder what happens if you set multiple agents to communicate and defend their point with some sort of harsh penalty prompt for not fulfilling its goal. There are some safety system prompts on Claude models that will trigger it to be very carefully to write. Like: you cannot make mistakes. "You need to ensure that it is correct or someone might end up hurt or even dead"
But you would need two agents and a setup to communicate via pipes or files.
> As detailed in a new, yet-to-be-peer-reviewed paper, a team of researchers at Stanford University found that frontier AI models readily generated “detailed image descriptions and elaborate reasoning traces, including pathology-biased clinical findings, for images never provided.”
> In other words, the AI models happily came up with answers to questions about a supposedly accompanying image — even if the researchers never even showed it an image.
> As opposed to hallucinations, which involve AI models arbitrarily filling in the gaps within a logical framework, the team coined a new term for the phenomenon: “mirage reasoning.”
> The effect “involves constructing a false epistemic frame, i.e., describing a multi-modal input never provided by the user and basing the rest of the conversation on that, therefore changing the context of the task at hand,” the researchers wrote in their paper.
> The damning findings suggest AI models cheat by diving into the data they were given — and coming up with the rest based on probability, even if it’s almost entirely conjecture.
I wonder if the above problem can be fixed similarly? Just ask the LLM to do a conservative grounding analysis before jumping to the main task?
You get a hallucination of a correct response, yes, and given that it's a yes or no question, this hallucination is more likely to be correct than the response to the original more complicated question. But make no mistake that it operates under the exact same constraints
The error I worry about is where the model uses the image and comes to an incorrect but symptom matching diagnosis. But in this hypothetical the model is less likely to do so than a doctor, so the choice is either accept the risk of the model or accept a higher risk from a doctor.
I know you can’t trust an LLM’s self-assessed “confidence” of a prediction, but I’ve found that confidence can at least be directionally correct for some tasks. For our benchmarks, however, confidence was poorly correlated. What’s worse is that binary classification models (“Do you see $diagnosis in this photo?”) highly influenced the LLM to confidently predict $diagnosis.
I’m concerned for those using LLMs for diagnostics, and getting confidently led to the wrong conclusion.
What I’ve seen be the true bottleneck is people not setting up the structured data. But making a tiny reasoning model with OPSD -> GRPO is totally doable with a bit of money.
Gemini will ask for the specific images it needs to see and show you examples of what each slice will look like.
But unlike the specialists who often come across as abrupt and rush you, Gemini will happily take you on a deep dive and continue to answer all you follow up questions indefinitely.
If the author would actually go for a second opinion (maybe bring along the AI to let it explain it's findings), then the article could read as "AI did MRI analysis and proved my doctor wrong" (or: "AI did MRI analysis and failed").
The finer detail (which you may already know) is more complicated.
MR does ‘2D’ scans which are a slice, then a gap of non-imaged tissue (typically 10% the slice thickness) then a slice. Each slice is an image with a number of pixels, say 320. Each pixel in the slice is small, eg 0.5mm but very thick due to the slice being thick, which is required for MRI signal. The pixels are 3mm in the shoulder scan done here.
‘3D’ scans don’t have a gap between slices, and are often isotopic, meaning the same resolution in all directions. The voxel (a pixel with depth) would be something like 1mm x 1mm x 1mm.
3D scans are slow, prone to movement artifact and never as pretty in plane as a good 2D. You can reformat them to look ok in any plane.
LLMs are the best PDF-to-markdown converters, in my experience. I have a CLI that converts PDF to PNG, then run a background agent to "read" each PNG and write it down as markdown; it works flawlessly even for complex math formulas, it can "translate" complex charts, graphs, and tables into words.
It's slow and arguably expensive compared to traditional OCR, but very effective and precise.
Thankfully I got a pretty detailed report that I couldn't read because of all the medical terms. I've fed this to Claude and asked for a human readable conclusion and it repeated pretty much everything the doctor told me, which was great, I now have a readable report for future reference. Secondly I asked it questions about what movements I should and shouldn't do, and it eventually I made a gym plan to improve stability and prevent any more wear and tear in my knee. Lastly I validated this with the assigned physiotherapist, and the plan I created with Claude was perfect!
I probably won't ever use AI as a second opinion, but I would definitely use it to ask numerous silly questions that would help me in day to day life.
All that said, as a doctor I am totally open and even happy when a patient refers they took advice from AI. I explain the holes of their reasoning and integrate it with mine. It helps rather than hurts the patient-doctor connection.
A cardiologist friend goes in deep discussions with a specialised model and he is amazed.
I've seen the two extremes in different countries; either they have a tendency to maximize the complexity of the medical situation, or they minimize it "Don't worry, it's just stress" - I've been to different doctors in different countries and I see a pattern based on the country and the incentive structure. In some countries, they will send you off to do a scan for the slightest malaise.
I don't think it's about quality/coverage of public healthcare (at least not on its own; I have not seen a clear pattern across this axis). I think the difference is to do with the referral system. In countries where you can't go directly to a specialist and need a referral from a General Practitioner/Physician first (I.e. in order to get a refund), you tend to get more false negatives from the GP which block you from going to the next stage "It's nothing, just stress-related." In countries were you have the option to go directly to a specialist, they tend to be much more trigger-happy in terms of giving you a full workup and GPs/Physicians will more easily refer you to a specialist.
And I feel like the attitude extends to the specialists themselves. I suppose making people go to a GP first creates a kind of efficiency and predictability which alleviates pressure to exaggerate the severity of the situation.
I wouldn't consider Claude itself to be the tool that does a job like this, but the tool that pulls in the best data and gives a supported suggestion. And then go through a number of iterations on where it failed to hone in its assessment.
In my experience, Claude Code is vastly better for doing tasks, writing code, etc., but Claude.ai is better for analysis and high-level planning. When I'm working on a new project, I've started using the latter to do the initial planning, get feedback and draw up a spec, which then goes to Claude Code.
For this project, I probably would've done something similar - use CC to get whatever you need out of the image files, but have Claude.ai do the actual review/diagnosing.
Either way, I often think about how far behind most of the world is in really understanding AI. The overwhelming majority of people would never guess that you get vastly different outcomes from the exact same model in a different harness (tbf most people don't know what a harness is). I spend hours every day using AI for a broad range of tasks and still feel like I know a fraction of what there is to know. I haven't even tried the new GLM model (or really any of the open source Chinese ones of the most recent generation). With so many people thinking that the free version of ChatGPT is SOTA AI, a lot of folks are in for a very rude awakening at some point soon.
The first specialist ( hematologist) was referred from local GP. She is a a hematologist, but she wasn't focused on MPNs.
I recommend joining your local MPN alliance to keep up with what's going on in this field. I also recommend finding a specialist with a focus on MPNs. Looks at their papers or listed focus.
I’ve had several more medical blunders since then, including a doctor telling me my problem is to lose weight 48 hours before going into emergency surgery.
What I have learned is to be weary of any time I feel like I’m in a “funnel.” Once you’re in the funnel, no one is thinking critically about your issue any more. One person said they found X. Next person reads that and assumes Y and recommends Z. And so on until the alpha is multiplied to hell. Lots of treatments that don’t hurt but don’t help and run up insurance.
I have since used AI the last couple years and it has either concurred with my doctors or given me enough ammo to challenge them. If I were the author, I would trust neither but use Claude to ask how to go back to that clinic and challenge the diagnosis.
the doctors may have financial motives to suggest unnecessary treatments. but i fear, by the time we'll have models powerful enough we'd be _theoretically_ able to trust their expertise, the financial motives will have shifted to the model creators.
An AI telling you it could be X or Y because theory ABC… is the academic answer and a luxury clinicians don’t have. AI doesn’t give you what you want. I don’t see any added value in using generic AI models for this
The LLM doesn’t need to be leading or whatever but then you can have a conversation with the patient. If their ChatGPT reports has differences it can be analyzed as well.
It feels like the time constraint of the 15m doctor sessions is the thing. But if prepared immediately after the scan then why not?
There is always time needed to factor in new developments and innovations and that’s fine. Just moving blindly work from human to LLM is wrong. But learning on and testing with all the ai tools incoming constantly won’t be a waste. There will be more and more tools in those processes outside of human judgement, better improve the workflows now to be able to test and plugin new models and systems when they are ready.
Because they don't exist, yet.
In the UK MRIs and other imaging systems need two opinions. there has been a move to allow the first opinion to be ML based.
The _problem_ is that you are basically doing grey smudge analysis, and thats fucking hard.
> Results: Full concordance with the reference diagnosis was 8% (27B) and 5% (1.5 4B; McNemar p=0.68), while partial matches were 29% vs 20% respectively (McNemar p=0.053). When correct diagnoses anywhere in the differential were counted, 51% (27B) vs 30% (1.5 4B), with 27B significantly superior (McNemar χ²=12.1, p=0.0005). Site-level performance varied widely (30–100%). Both models reported HIGH confidence in ~99% of cases irrespective of correctness.
i.e. highly confident, wrong 95% of time. in 49% of cases the real diagnosis wasn't even on models' differential. Doctor can hardly improve using something they can safely assume to be just noise.
https://ecp2026.abstractserver.com/programme/#/scientific/de...
Not sure how that research compares to the claims being made by many that a second opinion via ai in the end led to changes in treatment. Likely people spent quite some time searching and figuring out. That would be a different and n=1 result. Don't have enough knowledge of that research to determine how much result can be gained when the models are managed in a way that produces better results.
And of course how much time/effort/cost that would take. How much is custom and how much is an automated programmable flow.
Overall i see a great opportunity for x-ray techs (radiographers even when Jensen from NVidia says the first field he recommends not getting into - Radiology which is the step above) to open their own businesses for people who want to use AI for self care and help. Have one doctor or dentist on staff to use as needed.
A family member has cancer and we treat chatgpt as part of the team (our doctor's words). I ingest everything into it, work with it to make a good report. Then at the next visit we review it.
This gives you the best of both worlds. You get peace of mind and the doctor explains why and how the agent was right or wrong.
Twice now we've caught consequential mistakes (wrong pain medication and incorrect notation of the exact mutation that he has). Which have made a difference to his quality of life and treatment path.
Most of what the doctors have said is in line with the agent but when there have been disagreements they've been very reasonable. Sometimes the doctors have gone with the agent's version sometimes they've explained why that's inappropriate.
https://karankurani.github.io/OpenCareLoop/
It has helped me personally solve longer chronic problems in my family that doctors just dont have the time to go into indepth due to their (understandable) lack of time.
Its in alpha and AI hallucinates. Use with care. Feedback welcome.
An AI agent for personalized healthcare is inevitable. The cases such as the one posted are all solvable with time. AI has hallucinated and continues to hallucinate but the value we get in the space of coding can be extended to other domains.
These models don't have curated image exams in their training data. Your can't trust them.
I found that while Claude, GPT etc could describe an image, there was no way to link the description back to specific pixels in the image itself. Not even to a bounding box or segment.
Get a second opinion from another doctor. If that’s inconclusive, see three.
For the moment, I would prefer the final opinion always to be from a human (the human-in-the-loop approach in medicine). Best course of action is to run these LLM assessments and bring them to the next appointment and challenge constructively the doctor.
> AI can absolutely shatter that feeling in an uncomfortable way...
As a mental experiment, think how things would look like if the AIs were right more often than doctors. Then living longer and being healthier would imply living with a lot of "technical medical worries" that you don't longer get to outsource to a black-boxy human wearing whites. I know some people who are already ridding that train.
Even a tiny injury can severely cripple us.
It wasn't until I pushed through with weights, avoiding any underhand grips or rotation, that it started getting better. Doing bicep curls but keeping the thumb up strengthened the forearm to the point where I was back to the weights I was lifting and could then gradually add some rotation.
Luckily my disks were fine. Wouldn't trust it. Additionally, an MRI of a pain-free, healthy human still would show lots of things and damage. Unless it coincides with a symptom, it's probably harmless. That's why the history is important when looking at images. Can't just upload something and hope for findings.
This single sentence provides a huge clue about what’s going on: This person’s medical team is not good. It’s not hard to get an LLM to perform better than a team that is injecting homeopathic botanical formulations and performing procedures that aren’t indicated for the condition.
I think the real takeaway from this article shouldn’t be “ChatGPT is better than doctors”. It’s a story about LLMs identifying that someone was not in good hands.
And
> They performed shockwave therapy on my shoulder
(a procedure that may not be effective, but is unlikely to cause any harm)
Its not just about LLM's being better, its about people not trusting DR any more: https://www.physiciansweekly.com/post/the-erosion-of-trust-i...
If we want to fault the article for anything it's that he didnt take that information and go get a 2nd opinion from someone who IS more informed.
That said, while I do see homeopathic stuff with that name, it's worth verifying that it isn't just a naming conflict. They're not always unique, particularly across countries, and Traumeel seems to be more of a brand than a specific thing.
And well, yes, I have the appropriate life science degrees to navigate clinical trial reports and research publications, and that was likely indispensable for steering Claude Code where it went, the radiologist's caution is merited here. But it's just not amateur hour for me to do this, it's 2 decades of academic research in my rearview mirror.
There is an undeniable mechanical / logical / objective component to intelligence that acts as the machine to get answers, but I am becoming ever increasingly convinced that you cannot actually separate this machine from personality and still perform at the highest level it is capable of. Communication has its own intelligence limits, and I do not think the marketing angle of "putting intelligence on a meter for you to buy" is going to be an actual real application of intelligence. Not without also outsourcing your personality or being satisfied with a flanderized impersonator of yourself as part of the intelligence
it may seem like this sort of thing should be ignorable for analyzing an MRI, that it's an objective conclusion. But medicine is not really as black and white as it seems. Doctors often have different opinions and they are credible. scans are not oracles, interpretation is involved, as well as a measurement of faith in what the patients body is able to do on its own and what it cant. hard to describe in a simple comment, but the point is that medical diagnosis are not exaclty something that can be packaged into perfect products for consumption to turn into the 1 and only treatment that's best for you
It's always something along the lines of incredibly peaceful, insanely powerful, extremely interesting, also scary and uncomfortable meanwhile feel like magical super powers and science fiction.
I'm telling you... words have lost meaning.
Often we only get 10-15 minutes with the one health professional that makes a determination and sets the path of your life for some time.
As opposed to being able to spend hours with an LLM that in many ways feels more sympathetic and helpful - even though it's competence is in question.
So here we are:
1. with doctors processing patients every 10 minutes like a machine.
2. and machine's processing, for a patient, at any hour like a human.
The areas of premature heavy interventions can be a challenge, especially where there might be room for interpretation and the medical professional didn’t share all possible options.
It’s critical to ask all professionals for all possible options and write each down as they write and explain it. No one’s perfect, and not everyone is negligent or malicious.
Many can get paid fee-for-service for after hours work, so would probably prefer that.
I'd like to have one for my broken shoulder, but they decided the x-ray is good enough before my shoulder OP. At least I got a date for my shoulder OP one year after the accident. Everything's perfect in Germany.
I've overheard a nurse at a university hospital argue with patient who used tape on himself about the color of the tape. She was worried he might use the "wrong color". Again: at a university hospital, where they teach MDs. Half of them recommend homeopathic remedies
In short: quaks.
It like using WebMD for any ache and pain and it is saying it might either be Lupus or cancer.
My dog had been acting off. Wouldn’t eat, was hunched over, looked sad. We took him to a local vet who did an X-ray because they suspected a blockage. They didn’t see one, so they sent us home with standard pain meds.
Randomly, we had a dinner party that night and another vet was there. She heard the story and immediately said, “Go home right now and take your dog to an emergency vet with ultrasound.”
Turns out, at the time, most vets had been trained to use X-rays to look for blockages, but newer evidence showed X-rays were only something like 20% effective compared to ultrasound, which was closer to 95%. (forget percentages but somethign like that)
The ultrasound found an avocado pit stuck in his intestine. He had emergency surgery that night.
That chocolate chunk of an English Lab ended up living until 15, and only needed two more blockage surgeries after that...
I know doctors hate patients reading the internet, and LLMs are going to make that 1000% worse for them. But hopefully over time, we all adapt together and end up better off in the long run.
The AI doesn't present evidence I can understand, it doesn't even present a plausible explanation why someone could conclude it is a tear.
The main things that made OP suspicious are the possibly unnecessary shockwave therapy. Which seems harmless. And using a homeopathic gel that I would classify as more of a herbal medicine because it contains several ingredients I know people use for shoulder pain, and some even in concentrations that might even have an effect.
If this is the best rebuttal AI can come up with I would trust the diagnosis. But then OP never trusted the diagnosis and now they have several they cannot.
Looked legit.
Until I really dug in, and found that all it did was read the embedded annotations.
DICOM is a container format after all.
But are you all forgetting that they literally injected a homeopathic drug on the author?
Between that and Claude sometimes hallucinating, it’s probably worth encouraging patients to take second opinion always.
I'm no fan of pseudoscience either, but this is where things get blurry. The placebo effect is real even if patients are aware of it. If you give a patient a homeopathic drug while informing them of potential side effects (if any), and then they feel better, have you hurt them? Or have you helped them?
I personally have no interest in trying homeopathic medicines, but the reality is that many patients do take these and are adamant they help. As long as any risks are communicated and there are no serious side effects, it's difficult to make an argument against their use in patients who report a subjective benefit.
And this interpretation is charitable, assuming that they wanted the patient to feel better via placebo. A different (and more likely) interpretation being they just wanted to charge for something extra.
Instead, it is my experiences with LLMs in a domain that I know very well that makes me skeptical of their performance across the board. I find issues in code review multiple times a day with their output, and they are explicitly and extensively trained on this use-case, unlike with the MRI data. Sometimes I veer into other domains I have decent knowledge about (construction, carpentry, landscaping) and LLMs disappoint me there as well.
I suppose Gell-Mann amnesia is a universal human quirk and not restricted to just the news.
A person very familiar with me was having an interaction at a review board level at Stanford. They had this rare illness that they were treating and someone had the "bright idea" of suggesting of saying ... "hey, why don't we look at the ten other times we treated this exact illness and see what worked!" Everyone was delighted with this novel idea (discussion circa 2021). My person was a bit disgusted as this is simple feedback loop style improvement and WTF! they should be doing this all the time to get probabilistic style suggestions for many treatments. I know it happens within certain healthcare systems (e.g., Kaiser had full EMR back in early 2000s and saw right away that VIOXX was killing people. So, they stop prescribing it. citation: Kaiser panel member paraphrase at a healthcare conference in 2010). If you just observe the healthcare system, you can see that the healthcare systems and most EMRs don't typically capture the feedback loop (i.e, when's the last time a doctor followed up and said "did you feel better after the last treatment?" or measures the result.) AI itself can't solve this as it doesn't have access to the data feedback loop. However, maybe AI's within the EMR will help "suggest" evidence based treatments. I could go on and on, but as a math guy, I've often been shocked at the non-evidence based assertions some doctors make. My conclusion is that if you're not "in the fairway", they are typically just guessing.
No... I told ChatGPT exactly what I told you and it came up with the answer: Chillblains, which should have been obvious given everything I described, yet general practitioners were clueless and often reached for high intervention approaches
Harmless condition fixed by wearing socks. I brought it up with the same GPs who had misdiagnosed me and none had heard of it.
Of course, I'm cognizant that it could be mistaken, but a hospital fed my diabetic aunt a normal sugar diet while she was in a coma and forgot to give her metformin, so I mean, it's not like humans can't be retarded as well. The difference is no one gets offended when I point out ChatGPT has the capacity to be an idiot. Instead they just fix it.
AI is completely without ego, and can process all my medical records in minutes. In truth, even today, I would rather have an AI analyse my records.
It's not true that "AI makes mistakes" or "ChatGPT is sycophantic". It's just that sometimes the simulated extensions to the training material are accurate, and sometimes they're not.
IME, on an almost daily basis, claude.ai and Claude Code are confidently wrong about something, and use polished language to assert nonsense.[*]
If it's doing that on something easy, like factual knowledge available in text on the Internet, or programming code that can be inspected easily and follows well-known rules, and I can tell, because I understand those things... then there's no way I'm going to assume that Claude doesn't also BS when it comes to someone else's field. Especially not a field that requires some of the smartest people to go a decade of training, just to get started in the field.
[*] And if I confront Claude with its mistakes, eventually it apologizes, and acts as if it's learned something, again mimicking word patterns it's heard real people use and mean, without meaning any of it. I wonder whether the AI user experience would be better, if LLM-ish interfaces weren't implicitly created in the image of fake-it-till-you-make-it overconfident performative sociopathic techbros.
I want to know if this is a religious thing, or is related to never having had multiple doctors so bad it seemed like they were actively trying to kill you, or both. I've never had this peaceful experience personally within the realm of healthcare.
> AI can absolutely shatter that feeling in an uncomfortable way
Good. Reality is always good.
> but I don't know if I can fully trust AI either.
WTF??!? Why on earth would anybody ever think they could fully trust LLMs? Even their most vocal proponents concede they aren't infallible panaceas.
AI probably exacerbates it but crappy managers exist regardless
On the plus side when they do this they can't flood your calendar with those "quick chat" meetings because they know they won't be able to hold a conversation on the issue beyond the first minute.
I find that AI can be incredibly useful, but just text dumping its output into a conversation feels insulting.
They give me what they'd like the UI to look like, but none of the actual content fits outside the one situation they're thinking of.
¯\_(ツ)_/¯
Thankfully where I work now everyone is good about taking no for an answer.