IBM's Watson recommended 'unsafe and incorrect' cancer treatments
gizmodo.com
gizmodo.com
i.e. it's more like predictive text input than anything. If it makes it faster for you to diagnose, makes harder to diagnose stuff easier, and recommends the right treatment most of the time then it saves the doctor some time and energy.
The only question is whether it does that.
The model does not see a significant difference between the cat/ostrich, or cancer/cold, whereas we do; this implies that, when the model is wrong, it is likely to not just provide an incorrect treatment, but a catastrophically incorrect treatment.
Where the human sees a cat-like creature, and if not guess a cat, then something similar to a cat (4 legged, furry, etc), the ML model is willing to jump anywhere, in the worst case.
So its not just rate of misdiagnosis, but by how much as well.
If the doctor doesn't know how to quickly and cheaply verify the machine's guess, they'll just rubber stamp its recommendations.
People have a tendency to defer, and have a bias towards deferring to machines that seemingly behave as accurately as a calculator would with basic arithmetic.
A true crisis can arise if doctors can shift liability by claiming to have just followed Watson's results.
It increases the probability of (harmful) medical error. The "unsafe" is a big, big gotcha for regulatory approval. As someone correctly pointed out, this is relative to human error rates (can't view article because of registration).
A very expensive tool when healthcare systems are strapped for cash. Super unethical of IBM to try and extract profits here.
Do you realize they would have never developed Watson if there weren't profits involved?
I don't know about that. Most of the new construction in Connecticut seems to be building medical offices for Hartford Healthcare or other medical groups.
If you get complacent and just assume the computer knows what it's doing (because it usually does) this can and will end very badly.
The person whose Tesla drove into a concrete barrier at 65 mph earlier this year was perfectly capable of driving a car, but they mistakenly believed that the computer had it under control.
99% Invisible has a pair of podcast episodes on the subject (2015):
https://99percentinvisible.org/episode/children-of-the-magen...
https://99percentinvisible.org/episode/johnnycab-automation-...
Or maybe the tool should be limited to retrospective evaluation of doctors' decisions. Basically an automated peer review.
I'm all for automating driving and other such tasks once the computers are ready, but until we know they're at least most likely ready to do it, I want the ability to turn it off and do it myself.
(And frankly, I want that ability anyway because I genuinely love driving and it makes me sad to think someday I won't be able to do it.)
The fly-by-wire control system ordinarily prevents stalls, but it had disengaged due to an iced over sensor and was operating without stall protection. The plane stalled at 38,000 feet and fell into the ocean.
They should have had plenty of time to correct the stall, but one pilot was pulling back his control stick (the opposite of what they needed to do), and since the two sticks aren't physically linked the other pilot didn't know he was doing that.
It's one of the topics discussed in the podcast that I linked above.
We're still in pretty early days here, so I don't know if Watson's advice is at a point where that could be true. In the meantime, I do like the idea of having an unaided doctor and the computer program evaluate it independently, then afterward say "Ok, let's see what the computer thought" before coming to any final conclusions. Maybe it'd help avoid any dangerous recommendations. Or maybe people would say "Well the computer has a lot more data than me, it's probably right." I think that depends on whether the differences are errors similar to what a human would make, or if they're the "oh my god how did it even come up with this result" nonsensical error that are easily spotted by human review.
Two incidents demonstrate that:
1. https://en.wikipedia.org/wiki/Asiana_Airlines_Flight_214
2. https://en.wikipedia.org/wiki/Air_France_Flight_447
Pilots on both came from the civilian background. The difference between civilian training and combat training is that the latter is more trained to operate in conditions that involve instruments failure, distress, etc.
In the case of the Asiana flight, the captain chose a visual approach when he could have let the autopilot land it. And from there it's simple pilot error.
In the case of the air France flight, one of the pilots chose to ignore a stall warning and pull up instead of pushing down to gain speed. Sadly, in the case of no pilot input the plane actually would have ended up avoiding the stall on its own.
I of course agree that military pilots with real experience will likely perform better than civilian pilots, I just doubt autopilot had any part to play in either of those accidents, like the comment you replied to implies.
There is a talk[1] about this that is really popular among pilots, and I honestly think it seems increasingly relevant to the rest of the world as more and more things become computer automated in some fashion.
Oracle = make suggestion
Agent = Act on your behalf.
Waze might be starting off as an Oracle. but once you tie it into self driving cars, it'll become an agent acting on your behalf.
At a certain tipping point, creates really interesting questions. For example, if you re-direct 100% of traffic to less congested route, you end up creating traffic jams, so then you have to decide who goes on which route... maybe people who pay more get faster routing, or people deemed more important or busier by the algorithm...
"University of Texas says the project cost MD Anderson more than $62 million"
source: MD Anderson Benches IBM Watson In Setback For Artificial Intelligence In Medicine https://www.forbes.com/sites/matthewherper/2017/02/19/md-and...
As for your analogy, you got still tragically flawed predictive text technology for free. Nobody bought iPhones for their crappy predictive text capabilities. Imagine the horror if you actually had to pay for predictive text. As it is, it's a nice to have freebie that came along with your phone, which works.... sometimes.
Additionally, human-assisted A.I. is not the solution, it's a non-answer to the problem of creating systems that can think and perform at human levels of intelligence. It's okay to admit if we don't have the ability to make these things, but its disingenuous to believe that human involvement in helping computers get to the right answer is the right answer. Though yes, we need this right now to move things along where they otherwise might stand still.
ML performs very well for specific well defined tasks that have an obvious outcome and are highly narrow in scope.
Cancer is a disease that we can't even treat ourselves in many cases. It requires a great deal of creativity and critical thinking to reach solutions on a case by case basis. Who is so arrogant they thought this should be replaced by a bunch of overhyped software?
Maybe it's counterintuitive, but it makes a lot more sense to use ML for cancer than say a broken arm. There are so many systems interacting in cancer that affect its progression that humans really are at there limits in trying to understand them.
And as others have said, no one is letting Baymax loose in the oncology ward and firing all the doctors. This is just one more tool in a doc's tool belt -- and far from the only that will give misleading results.
Bacterial infections also have a ridiculous amount of systems involved and yet that was figured out by humans.
I don't see how ML helps with cancer at all right now. The problem isn't the amount of data. It's the quality of it.
OK. Back in the real world, successful AI companies build useful software products that inform human decision making but do not directly, unilaterally, unintermediatedly result in real-world kinetic action.
Why not look through the papers each company has released though and decide for yourself? https://ai.google/research/teams/applied-science/quantum-ai/ https://www.ibm.com/blogs/research/category/quantcomp/
That said, I don't think IBM Watson has published much top-tier research.
It’s weird because IBM spends a lot on research (6 billion when I was there) but they are always trying to get marketable things out of said research. They still had chip fabs when I was there so they had people working on chemistry physics and math for chip things. They had chess playing machines (deep blue) and a backgammon one using ai.
They hired a lot of phds who were for the most part very self motivated. Very smart people. I’m not sure if they are still getting the best and brightest but for a lot the IBM letters are a draw.
a biopsy from your peripheral plasma, opposed to solid tumor. A big issue in cancer medicine is it's super fucking hard to get any kind of noninvasive measurement. Typically it's done with surgery or with CAT scans, which have extremely low precision.
Research in the last few years has been pointing to the idea that we can detect tumor derived DNA fragments in plasma. The challenge being, healthy dna makes up 99.9% of it, this means that currently the methods only work when the tumor burden (metastatic for instance) is high. Not ideal for early detection or treatment monitoring. But if the computational tools improve (rn none assume such a low mixture), you could see sensitivity and precision increase to the point that it's useful for predicting therapeutic outcomes and for early detection.
i guess no, they were not even relevant, which is why doctors were not recommending them
It can recommend unconsidered, more effective treatments just as readily as it can recommend unconsidered, dangerous ones — particularly if the system was as extensively trained using MSK's physicians' preferred treatments, instead of (or even additionally to) actual clinical data as The Fine Article suggests.
Or is this just a plain old failure?
It was never more than a cool Jeopardy robot and some sweet Bob Dylan ads
Ever? I don't see why not. In our lifetime? Probably not and especially not with current technology.