GPT-4V(ision) Unsuitable for Clinical Care and Education: An Evaluation
arxiv.org
arxiv.org
> AI company releases generalist model for testing/experimentation
> Users unwisely treat it like a universal oracle and give it tasks far outside its training domain
> It doesn't perform well
> People are shocked and warn about the "dangers of AI"
This happens every time. Why can't we treat AI tools like they actually are: interesting demonstrations of emergent intelligent properties that are a few versions away from production-ready capabilities?
I wrote a rant a few months back about the greatest threat to generative AI is people using it poorly: https://minimaxir.com/2023/10/ai-sturgeons-law/
That, and these companies have a substantial financial interest in pushing the omniscience/omnipotence narrative. OpenAI trying to encourage responsible AI usage is like Phillip Morris trying to encourage responsible tobacco use. Fundamental conflict of interest.
As for the future I am certain LLMs will become more efficient in terms of resource consumption and easier to train, but I am not so certain that they're going to fundamentally solve the problems that LLMs have now.
Try to train one to tell you what kind of shoes a person is wearing in an image and it will likely "short circuit" and conclude that a person with fancy clothes is wearing fancy shoes (true much more often than not) even if you can't see their shoes at all. (Is a person wearing a sports jersey and holding a basketball on a basketball court necessarily wearing sneakers?) This is one of those cases where bias looks like it is giving better performance but show a system like that a basketball player wearing combat boots and it will look stupid. So much of the apparent high performance LLMs comes out of this bias and I'm not sure the problem can really be fixed.
Problems like that have been bugging me for a while, if you had some Bayesian model you could adjust the priors to make the statistics match a particular set of cases and it would be nice if you could do the same with a wide range of machine learning approaches. For instance you might find that 90% of the cases are "easy" cases that seem like a waste to include in the training data, keeping there gives the model the right idea about the statistics but may burn up the learning capacity of the model such that it can't really learn from the other 10%.
I talked w/ some contractors who made intelligence models for three-letter agencies and they told me all about ways for dealing with that come down to building multiple-staged models where you build a model that separates the hard cases from the easy cases with specialized training sets for each one. It's one of those things that some people have forgotten in the Research Triangle Park area but Silicon Valley never knew.
Here's the shoe-identification experiment with GPT-4 and it's not confused by a suit jacket:
It's also not confused by a tuxedo:
Also tried the same examples with Google Gemini 1.5 Pro and it doesn't hallucinate either.
He said the same thing about perceptrons in general. When it comes to bunk, Minsky was... let's just say he was a subject matter expert.
He caused a lot of people to waste a lot of time. I've got your XOR right here, Marvin...
Just add another model asking it if there are shoes visible or not.
You probably need to segment out the feet and then have a model that just looks at the feet. Just throwing out images without feet isn't going to tell the system that it is only supported to look at the feet. And even if you do that, there is also the inconvenient truth that a system that looks at everything else could still beat a "proper" model because there are going to be cases where the view of the feet is not so good and exploiting the bias is going to help.
This weekend I might find the time to mate my image sorter to a CLIP-then-classical-ML classifier and I'll see how good it does. I expect it to do really well on "indoors vs outdoors" but would not expect it to do well on shoes (other than by cheating) unless I put a lot of effort into something better.
When we see a thing with coherent writing about any subject we're not experts in, even when we notice the huge and dramatic "wet pavements cause rain" level flaws when it's writing about our own speciality, we forget all those examples of flaws the moment the page changes and we revert once more to thinking it is a font of wisdom.
We've been doing this with newspapers for a century or two before Michael Crichton coined the Gell-Mann Amnesia effect.
A quick Google search found this paper, here's a quote:
"Our evaluation shows that GPT-4V excels in understanding medical images and is able to generate high-quality radiology reports and effectively answer questions about medical images. Meanwhile, it is found that its performance for medical visual grounding needs to be substantially improved. In addition, we observe the discrepancy between the evaluation outcome from quantitative analysis and that from human evaluation. This discrepancy suggests the limitations of conventional metrics in assessing the performance of large language models like GPT-4V and the necessity of developing new metrics for automatic quantitative analysis."
https://arxiv.org/html/2310.20381v5
For sure there's some waffling at the end, but many people will come away with the feeling that this is something GPT-4V can do.
If an ai is able to record an operation with 9x% accuracy but you save humans, the insurance might just accept this.
Nonetheless the chance that ai will be able to continually become better and better at it, is very high.
Our society will switch. Instead of rewriting software or updating software we will fine-tune models and add more examples.
Because this is actually sustainable (you can reuse the old data) this will win at the end.
The only thing changing in the future is model architecture and training data will only be added.
"It is difficult to get a man to understand something when his salary depends on his not understanding it."
1. Because you've got one or more of the below spinning it into either a butterfly to chase or a product to buy:
- 'Research Groups' e.x. Gartner
- Startups with an 'AI' product
- Startups that add an 'AI' feature
- OpenAI [0]
2. I'm currently working on a theory that a reasonable portion of population in certain circles is viewing ChatGPT and it's ilk as the perfect way to mask their long COVID symptoms and thus embracing blindly. [1]
[0] - The level of hyuperis in some articles about ChatGPT3 reminded me a little too much of the fake viral news around the launch of Pokemon Go, adjusted for fake viral news producers improving quality of tactics. Especially because it flares up when -they- do things... but others?
[1] - Whoever needs to read this probably won't, but; I know when you had ChatGPT wrote the JIRA requirements and more importantly I know when you didn't sanity check what it spit out.
> A friend sent me MRI brain scan results and I put it through Claude.
> No other AI would provide a diagnosis, Claude did.
> Claude found an aggressive tumour.
> The radiologist report came back clean.
> I annoyed the radiologists until they re-checked. They did so with 3 radiologists and their own AI. Came back clean, so looks like Claude was wrong.
> But looks how convincing Claude sounds! We're still early...
And I don't know any, but there is probably already tools to help radiologists, what they call their "own AI".
https://blog.jetbrains.com/blog/2023/10/16/ai-graphics-at-je...
For whatever reason Hugo's docu is weird to get into while chatgpt is shockingly good in telling me what I actually look for
Looking at an ECG as a layperson it's a problem that seems easy if you know about some tools in your math toolbox, but it's deceptively hard and a false negative might mean death for the patient. So, I'm not going to trust a generic vision transformer model to this task, and until I see overwhelming evidence I won't trust a specifically trained model for this task.
I'd trust a study like this a little more if the human evaluators were presented with the output of GPT-4 mixed together with the output from human experts, such that they didn't know if the explanation they were evaluating came from a human or an LLM.
This would reduce the risk that participants in the study, consciously or subconsciously, marked down AI results because they knew them to be AI results.
But until it comes to fruition, I think it's largely a waste for people to spend time studying the viability of general models for medical tasks.
> All images were evaluated by two senior surgical residents (K.R.A, H.S.) and a board-certified internal medicine physician (A.T.). ECGs and clinical photos of dermatologic conditions were additionally evaluated by a board-certified cardiac electrophysiologist (A.H.) and dermatologist (A.C.), respectively
In the paper the tasks are only completed by GPT-4V. For a valid scientific investigation, there should be a control set completed by e.g. qualified doctors. When the panel of experts does their evaluation, they should rate both sets of responses so that the difference in score can be compared in the paper.
It assumes that "here are 5 doctors which are always correct". Then measures GPT's correctness against them.
Clearly we are not there yet, but within certain conditions we are getting close. If we enable nurses and doctors to be more productive with these tools, we absolutely should explore that. Keep a doctor in the loop as AI is still error prone, but you can line things up in a way such that doctors make the decision with AI gathered information.
Run research on models that are actually trained to solve these issues, that is relevant research
Just as an example, Viz.ai applied[1] and received FDA clearance/approval for their model in hospitals.
Have OpenAI ever submitted a request for use of GPT-4V? Whats next, try autopilot driving with GPT-4V?
[1] https://www.viz.ai/news/viz-ai-receives-fda-510k-clearance-f...