Claude thinks it's a tumor, radiologists disagree
twitter.com
twitter.com
2) I am not sure its surprising that a general purpose model is not great at reading MRI scans.
G.Hinton, 2016
> The error rate produced by a deep learning based methods dropped below that achieved by human observers in 2015 for the first time
Essentially we just need good implementations at this point. We're there, just not widely deployed.
I do not believe in the infallibility of medical professionals. And I have seen AI totally get it wrong every day. But, I am of the opinion that we are not very far from the point where AI will be better at diagnosing issues than an average doctor. That isn't a high bar. An average doctor isn't very good.
I think AI is almost there already for an average programmer. An average programmer isn't very good either.
Perhaps using general purpose LLMs for safety critical use cases without humans isn't a good idea. This is why we have a 'human in the loop' giving the feedback to resolve the trust issue that LLMs lack.
Maybe we should trust LLMs like Claude to pilot a plane without any human pilots on board and raise VC money on the idea on replacing all pilots since we keep having these conversations on LLMs being compared with humans and countless VCs screaming over robots replacing humans in all occupations and 'accelerating everything' which could not be more wrong.
Chat-GPT and all the LLMs are doing essentially the same "predict the next work in a sentence" to sound as plausibly real as possible, which is essentially the same gimmick, only we call it hallucination due to VC funded AI firms need to pretend LLMs understand anything.
ChatGPT does specifically disallow medical advice and generally it is RLHFed not to provide it: https://openai.com/policies/usage-policies