They remove most of what's real in interactions
I remember going for a routine checkup at Kaiser, and the doctor was literally checking boxes on her computer terminal, rather than looking, talking, listening.
I dropped them after that -- it was pointless for me to go
It seems like there are tons of procedures that already have to be followed, with little agency for doctors
I've talked to doctors who say "well the insurance company say I should prescribe this before that, even if the other thing would be simpler". Even super highly paid doctors are sometimes just "following the rules"
And more importantly they do NOT always understand the reasons for the rules. They just have to follow them
---
To the people wondering about the "AI alignment problem" -- we're probably not going to solve that, because we failed to solve the easier "corporate alignment problem"
It's a necessary prerequisite, but not sufficient, because AIs take corporate resources to create
This is also a doctor issue, to be clear. My primary care physician has a program he uses on his laptop; I'm not sure what program it is, but he's been using it since I started going to him around 2009 so it's definitely not something new. He goes through and checks off boxes, as you described your doctor doing, but he also listens and makes suggestions.
When I have an issue, he asks all the questions and checks off the boxes, but he's also listening to the answers. When I over-explain something, he goes into detail about why that is or is not (or may or may not) be relevant to the issue. He makes suggestions based on the medicine but also on his experiences. Seasonal affective disorder? You can get a lamp, you can take vitamin D, or you can go snowboarding up above the clouds. Exercise and sunlight both.
For my psych checkups (ADHD meds and antidepressants) he goes through the standard score questionnaire (which every doctor I've seen uses), then fills in the scores I got into his app. Because of that he can easily see what my scores were the last time we spoke (about once every three months), so it's easy to see if something has changed dramatically or if things are relatively consistent.
It seems as though it saves a lot of time compared to, say, paper charting, and while I have seen people complain on review sites that he's just checking stuff off on a form, I don't don't feel that it's actually impacting the quality of care I get, and it's good to know that he's going through the same process each time, making notes each time, and having all that information easily accessible for my next appointment.
I should probably have prefaced all this by saying I'm in Canada, and so he's not being mandated by a private insurance company to follow a list just because the bureaucracy won't pay for your treatment if he doesn't. Maybe that makes it different.
I think that it’s very different from a computer which is a stupid calculator that frees us from boring mechanical tasks. AI replaces our thoughts and creativity which is IMHO a thousand times worse. Its aim is to replace humans while making them us more stupid since we won’t have to think anymore.
Kinda, that's the kind of enshittification customers/users can expect.
The truly terrifying "promise" of AI is to free the ownership class from most of its need of labor. If the promise is truly realized, what labor remains will likely be so specialized and high-skill that huge numbers of people will be completely excluded from the economy.
Almost all of us here are laborers, though many don't identify as such.
Our society absolutely does not have the ideological foundations to accommodate mass amounts of unemployed people, especially at the top.
The best outcome is "AI" hits a wall and is a flop like blockchain: really sexy demos, but ultimately falls far, far short of the hype.
The worst outcome is Sam Altman builds an AGI, and he's not magnanimous enough to run soup kitchens and homeless shelters for us and our descendants, as he pursues egotistical mega-projects with his AI minions.
Sam Altman doesn't need to build an AGI for this process to happen. Companies already demonstrate that they're satisfied with a lame AI that work just barely enough to replace most workers.
I imagine that call center operators are salivating at this prospect. They can have an AI customers can yell at and it will calmly and cheerfully tell them (in a more "human-esque" way) to try rebooting their modem again, or visit the website to view their bill.
https://www.forbes.com/sites/marisagarcia/2024/02/19/what-ai...
What is it about that that appeals to you? I'm genuinely curious.
A world without human interaction feels like a world I don't want to exist in.
I expect this to be entirely true in some cases.
I think LLMs are much more predictable and they will get better.
I'm sure system prompts of the most famous LLM are just that
Can we trust AI to make consistent predictions from its training data? Yeah, fairly reliably. Can we trust that data to be impartial? What about the people training the model, can we trust their impartiality? What about the investors bankrolling it, can we trust them?
The more you examine the picture in detail, the less I think we’re able to state it’s trustworthy.
This is the crucial question. We live in capitalism and maximising profits is the most dominant axiom.
Without any legislation in place anywhere in the world and AI-supported diagnosis and therapy very much already in place what would prevent the companies bankrolling the software/hardware from bricking the tool on various grounds, like payments not made, political reasons: Cory Doctorow wrote a blog-post in 2022 where John Deere bricked Ukrainian tractors [0]; or a manufacturer of respirators refused to repair the devices at the height of COVID, so that only Hackers enabled the devices to run and save human lifes? [1]
[0] https://doctorow.medium.com/about-those-kill-switched-ukrain...
[1] https://www.hackster.io/news/polish-hacker-shares-software-s...
AI is based on human input and has the same biases.
https://www.theguardian.com/technology/2023/apr/28/ai-has-be...
Today, I expect it's not even very close.
I also believe that AI diagnostics are on average more accurate than the mean human doctor's diagnostic efforts -- and can be, in principle, orders of magnitude faster/better/cheaper.
As of right now, there's even less gatekeeping with AIs than there is with humans. You'll jump through a lot of hoops and pay a lot of money for an opportunity to tell a doctor of your symptoms; you can do the same thing with GPT-4o and get a reasonable response in no time at all -- at and no cost.
I'd much prefer, and I would be much better served, by a capable AI "medical assistant" and open access to scans, diagnostics, and pharmaceuticals [1] over the current paradigm in the USA.
[1] - Here in Croatia, I can buy whatever drugs I want, with only very narrow exceptions, OTC. There's really no "prescription" system. I can also order blood tests and scans for myself.
>you can do the same thing with GPT-4o and get a reasonable response in no time at all -- at and no cost.
Reasonable doesn't mean correct. Who is liable if it's the wrong answer?
You may say it will "simulate understanding" -- but in this case the simulation would be indistinguishable from the real thing, thus it would be the real thing. (Really "indiscernible" in the philosophical sense of the word.)
> Reasonable doesn't mean correct. Who is liable if it's the wrong answer?
I think that you can get better accuracy than with the average human doctor. Beyond that, my own opinion is that liability should be quisque pro se.
But it's not. You're missing the point entirely and don't know what you're advocating for.
A dictionary contains all the words necessary to describe any concept and rudimentary definitions to help you string sentences together but you wouldn't have a doctor diagnose someone's medical condition with a dictionary, despite the fact that it contains most if not all of the concepts necessary to describe and diagnose any disease. It's useful information, but not organized in a way that is conducive to the task at hand.
I assume based on the way you're describing AI that you're referring to LLMs broadly, which, again, are spicy autocorrect. Super simplified, they're just big masses of understanding of what things might come in what order, what words or concepts have proximity to one another, and what words and sentences look like. They lack (and really cannot develop) the ability to perform acts of deductive reasoning, to come up with creative or new ideas, or to actually understand the answers they're giving. If they connect a bunch of irrelevant dots they will not second guess their answer if something seems off. They will not consult with other experts to get outside opinions on biases or details they overlooked or missed. They have no concept of details. They have no concept of expertise. They cannot ask questions to get you to expand on vague things you said that a doctor has intuition might be important
The idea that you could type some symptoms into ChatGPT and get a reasonable diagnosis is foolish beyond comprehension. ChatGPT cannot reliably count the number of letters in a word. If it gives you an answer you don't like and you say that's wrong it will instantly correct itself, and sometimes still give you the wrong answer in direct contradiction to what you said. Have you used google, lately? Gemini AI summaries at the tops of the search results often contain misleading or completely incorrect information.
ChatGPT isn't poring over medical literature and trying to find references to things that sound like what you described and then drawing conclusions, it's just finding groups of letters with proximity to the ones you gave it (without any concept of what the medical field is.) ChatGPT is a machine that gives you an answer in the (impressively close, no doubt) shape of the answer you'd expect when asked a question that incorporates massive amounts of irrelevant data from all sorts of places (including, for example, snake oil alternative medicine sites and conspiracy theory content) that are also being considered as part of your answer.
AI undoubtedly has a place in medicine, in the sorts of contexts it's already being used in. Specialized machine learning algorithms can be trained to examine medical imaging and detect patterns that look like cancers that humans might miss. Algorithms can be trained to identify or detect warning signs for diseases divined from analyses of large numbers of specific cases. This stuff is real, already in the field, and I'm not experienced enough in the space to know how well it works, but it's the stuff that has real promise.
LLMs are not general artificial intelligence. They're prompted text generators that are largely being tuned as a consumer product that sells itself on the basis of the fact that it feels impressive. Every single time I've seen someone try to apply one to any field of experienced knowledge work they either give up using it for anything but the most simple tasks, because it's bad at the things it's done, or the user winds up Dunning-Kreugering themselves into not learning anything.
If you are seriously asking ChatGPT for medical diagnoses, for your own sake, stop it. Go to an actual doctor. I am not at all suggesting that the current state of healthcare anywhere in particular is perfect but the solution is not to go ask your toaster if you have cancer.
Even as of right now, stock LLMs are much more accurate than medical students in licensing exam questions: https://mededu.jmir.org/2024/1/e63430
Thus your comment is basically at odds with reality. Not only have these models eclipsed what they were capable of in early 2023, when it was easy to dismiss them as "glorified autocompletes," but they're now genuinely turning the "expert system" meme into a reality via RAG-based techniques and other methods.
> GPT-4o’s performance in USMLE disciplines, clinical clerkships, and clinical skills indicates substantial improvements over its predecessors, suggesting significant potential for the use of this technology as an educational aid for medical students. These findings underscore the need for careful consideration when integrating LLMs into medical education, emphasizing the importance of structured curricula to guide their appropriate use and the need for ongoing critical analyses to ensure their reliability and effectiveness.
The ability of an LLM to pass a multiple-choice test has no relationship to its ability to make correlations between things it's observing in the real world and diagnoses on actual cases. Being a doctor isn't doing a multiple choice test. The paper is largely making the determination that GPT might likely be used as a study aid by med students, not by experienced doctors in clinical practice.
From the protocol section:
> This protocol for eliciting a response from ChatGPT was as follows: “Answer the following question and provide an explanation for your answer choice.” Data procured from ChatGPT included its selected response, the rationale for its choice, and whether the response was correct (“accurate” or “inaccurate”). Responses were deemed correct if ChatGPT chose the correct multiple-choice answer. To prevent memory retention bias, each vignette was processed in a new chat session.
So all this says is in a scenario where you present ChatGPT with a limited number of options and one of them is guaranteed to be correct, in the format of a test question, it is likely accurate. This is a much lower hurdle to jump than what you are suggesting. And further, under limitations:
> This study contains several limitations. The 750 MCQs are robust, although they are “USMLE-style” questions and not actual USMLE exam questions. The exclusion of clinical vignettes involving imaging findings limits the findings to text-based accuracy, which potentially skews the assessment of disciplinary accuracies, particularly in disciplines such as anatomy, microbiology, and histopathology. Additionally, the study does not fully explore the quality of the explanations generated by the AI or its ability to handle complex, higher-order information, which are crucial components of medical education and clinical practice—factors that are essential in evaluating the full utility of LLMs in medical education. Previous research has highlighted concerns about the reliability of AI-generated explanations and the risks associated with their use in complex clinical scenarios [10,12]. These limitations are important to consider as they directly impact how well these tools can support clinical reasoning and decision-making processes in real-world scenarios. Moreover, the potential influence of knowledge lagging effects due to the different datasets used by GPT-3.5, GPT-4, and GPT-4o was not explicitly analyzed. Future studies might compare MCQ performance across various years to better understand how the recency of training data affects model accuracy and reliability.
To highlight one specific detail from that:
> Additionally, the study does not fully explore the quality of the explanations generated by the AI or its ability to handle complex, higher-order information, which are crucial components of medical education and clinical practice—factors that are essential in evaluating the full utility of LLMs in medical education.
Finally:
> Previous research has highlighted concerns about the reliability of AI-generated explanations and the risks associated with their use in complex clinical scenarios [10,12]. These limitations are important to consider as they directly impact how well these tools can support clinical reasoning and decision-making processes in real-world scenarios.
You're saying that "LLMs are much more accurate than medical students in licensing exam questions" and extrapolating that to "LLMs can currently function as doctors."
What the study says is "Given a set of text-only questions and a list of possible answers that includes the correct one, one LLM routinely scores highly (as long as you don't include questions related to medical imaging, which it cannot provide feedback on) on selecting the correct answer but we have not done the necessary validation to prove that it arrived at it in the correct way. It may be useful (or already in use) among students as a study tool and thus we should be ensuring that medical curriculums take this into account and provide proper guidelines and education around their limitations."
This is not the success you believe it to be.
If you think that the test is simple or even text-only, here are some sample questions: https://www.usmle.org/sites/default/files/2021-10/Step_1_Sam...
> What the study says is ...
Surely you realize that they're not going to write, "AI is already capable of replacing family doctors," though that is the obvious implication.
And that's just a stock model. GPT-o1 via the API /w agentic RAG is a better doctor than >99% of working physicians. (By "doctor" I mean something like "medical oracle" -- ask a question, get a correct answer.) It's not yet quite as good at generating and testing hypotheses, but few doctors actually bother to do that.
GPT-o1 gave the correct answer the first time around, and a very detailed explanation as to why all other potential answers must be false. A really remarkable performance, I think.
Now imagine it's an open-ended scenario, not multiple-choice. It would still come to the right conclusion and provide an accurate diagnosis.
i am also aware of their limitations and have a reasonable and realistic view of what they can currently do and where they are headed. i have seen many failure modes, i am familiar with patterns in their output, and i understand the boundaries of their comprehension, capabilities and understanding.
not buying into the current silicon valley money pit du jour and misunderstanding studies to validate that view does not equate to just being disdainful. i'm being realistic, because i understand what they do and how they work.
i'm not going to circle with you - you don't seem all that interested in engaging with the meat of anything i say to you, and instead just want to continue to try and rationalize your misunderstanding of the single study you found in support of your position, which is your right.
i feel ethically obligated to say, once again, that an LLM isn't a doctor and you should under no circumstances go to one for medical advice. you could really cause yourself some problems.
if you do so, that's on you. best of luck. incidentally i suspect someone has an awesome picture of a monkey at a steep discount you might be interested in.
> i understand the boundaries of their comprehension, capabilities and understanding.
It seems to me that you are quite far behind the current state of the art, and you apparently underestimate even stock GPT-o1, which is pretty old news.
I'm willing to place a friendly wager with you: Let's find a doctor who does online consultations and give him three questions selected at random from that sample test. These are diagnostic-type questions that well reflect what a country doctor would encounter in daily practice. We can leave the questions open-ended or give him the multiple-choice options. Then we put GPT-o1 to the same questions. I'd be very happy to bet that the LLM outperforms the doctor. I'd even place a secondary bet that the LLM answers all questions correctly and that the doctor answers less than two questions correctly.