IBM Watson Overpromised and Underdelivered on AI Health Care
spectrum.ieee.org
spectrum.ieee.org
- if you're a generally competent, learned person
- and the new product feels like a giant leap from anything that exists already
- it's likely not all it claims to be.
If anyone is promising a 20% or greater improvement over the current leaders, I immediately file it away as a scam until I see extraordinary proof.
Other than some very immature industries, even 20% leaps don't happen without earth shattering once-in-a-generation breakthroughs.
Some call me a pessimist, and occasionally I'm unfairly dismissive of new technologies. But for every one I get wrong, I get 99 others correct.
Mature and/or competitive.
20%+ without a breakthrough can absolutely happen in niche industries that have historically lacked competition.
What I think made people miss it is no one wanted to refactor their code to run efficiently on GPUs. It wasn't that hard to do so in a lot of cases but it was work.
And all this set the stage for deep learning. It was the opening act if you will...
CUDA offered a way to access a limited set of operations on special-purpose hardware with a rewrite of critical sections. A pretty far cry from 'boosting legacy code.'
Ha, this is just silly. Ok, so what percent were IBM promising? How do you ever reduce anything complicated into a percent change.
Sell consulting deployments that hope to fund development of promised features of the tech.
Consulting is a bit like used car selling. If you're not playing close to the line, then someone else is.
Which, okay. Is what it is. But every customer should be aware this is what IBM et al.'s business model is, unless you're simply buying bodies.
Maybe I'm just naive, but I really think that strand still exists in IBM's corporate DNA. They still see this as solvable given sufficient time and effort.
When you realize how much positioning is required for large scale sales, you can start to understand why the same company might not have the feedback loops in place to build a product.
100% correct. Had IBM marketed Watson for what it actually is - an enterprise version of Amazon Alexa or Google Assistant (which admittedly didn't even exist at the time) - they'd be in a completely different spot. Instead we got the Bob Dylan commercials (which under the covers was just speech to text, a hardcoded dialog tree, and text to speech) and all the "Watson is curing cancer" nonsense (which under the covers was just a rules engine sitting on top of a search engine intended to provide customized treatment plans). The actual technology works, but it's an NLP platform for building chatbots and search engines (think intent classifier, relevancy ranker, named entity/semantic relationship/sentiment annotators) not the AGI that it was marketed as.
* Whatever the quality of the technology (which I personally never saw as that compelling) was wrapped up in terribly written research code, making it practically impossible to setup and use.
* The Jeopardy demo was made possible by the existence of a marked-up source of general knowledge (Wikipedia), a ready-made bank of questions and answers from past shows (j-archive.org), and the fact that practically anyone had the ability to curate more Q&A pairs. This is almost totally different than the medical use case where the knowledge is wrapped up in proprietary textbooks and papers and the only people able to curate training data are medical professionals.
https://www.quora.com/Why-isn%E2%80%99t-machine-learning-mor...
Skip to Jae Won Joe's answer to see a case study with a patient showing the interactive process. Then, especially look at ending questions about missing or deceptive data. Seems like their good diagnoses comes from a combination of domain data and expertly reading people in front of them. Machines suck at the part I emphasized. The data sets might not reflect a lot of stuff like that, too. Who knows.
It would take a paradigm shift before AI became useful. Going to watson.com and getting a prescription for antibiotics would be useful and technically feasible. So would augmenting lower trained people to be able to deliver more care. Neither are possible legally and it's not exactly a field you can get away with disregarding the law.
* Procure patient medical files. This rabbit hole includes things like HIPPA compliance and sufficiently anonymizing the files. Also, hospitals consider 3 years of information on 100 patients to be a lot of data.
* Have trained medical professionals annotate the files. Good luck convincing doctors to do this. Also good luck in determining what to do if the handful of doctors you're granted a few hours with a week can't agree on annotations.
* Actually produce a model that is reasonably good at pulling information out of patient charts,
Even after all the above, you still only have a solution that is giving you structured data, which is something that other EMRs can provide w/o the problems of accuracy & completeness of NLP process. There is still the process of diagnosis and treatment and also the fact that this is all a loop ('active learning').
All hype and marketing built on open source tools, all while AWS, Google, and Azure build out bigger cloud offerings. As OP mentioned, they used to be a great company.
If your answer is mission critical, probabilistic ML isn't advanced, explainable, or reliable enough for your problem.
If your answer is of the great-to-be-right, meh-to-be-wrong sort (e.g. ranking movie recommendations), then you can and should go nuts with ML.
And if someone really wants to do an ML project on the former, do everything you can to transform it into the latter.
It seems like a problem when classifying user-provided images (eg. identifying obscene images on a social network) but not so relevant when you own the sensors.
But some projects as presented don't have a clear refusal option.
ML would go a long way if the Hippocratic Oath (or derivative thereof) were taken & adhered to by the industry.
Seriously, I used IBM Watson on a consulting project a little over two years ago and I was disappointed. To be fair I should take another look. IBM has bought some good companies whose work is exposed in BlueBiz and other web services and the permanent free tier levels they give away rival the free tier levels from GCP.
I think AutoML and auto data science web hosted offerings will be a huge market with right now Google, Amazon, and Microsoft leading the way. If I worked at IBM in a position of power, I would work hard to create a great developer experience for simplifying the use of machine learning, NLP, etc.
With deep learning it took the transition seriously internally as well: there's a huge push to get all models to use Tensorflow - even if the transition itself is painful - and this helps the infrastructure providers as well to gather feedback.
With IBM it all started as a very successful research project, but instead of working more on the Watson system, it decided to rename dumber projects to be under the Watson umbrella as well thereby giving the impression that it's the same system that won in Jeopardy.
It wasn’t bad at all, we didn’t cure cancer but it provided a not so awful way to do some basic ML workloads. The other problem is that IBM pitches Watson to anyone, and many of the customers lack the clue to utilize it.
Anyone surprised by this?
For instance, they developed software for the Apollo mission and the Space Shuttle. The IBM/360. The IBM PC. AS/400.
Do you have a more recent example? something that happened several years after Louis Gerstner first assumed leadership.
Admittedly though given their pretty large size I can't name much else.
Exactly. Even if 20,000 people worked on those supercomputers, mainframes and Watson, what do the other 350,000 employees work on? Consulting, and it's been that way since Gerstner.
I hope that's not the case for their quantum computer
The only explanation I've been able to come up with is that when it's all human processed, Joe 2nd-link-in-the-chain just deals with all the inconsistencies as best he can to get his job done, and never reports issues up.
It generally seems like (a) the suggestions for fixes are impractical to implement (overly detrimental effect on counterparty), (b) Joe isn't empowered organizationally to suggest fixes that will be implemented, or (c) Joe doesn't have access to the IT tools to implement fixes himself.
After using that, I realized that Watson was mostly a gimmick.
https://twitter.com/chaoticmass/status/677708888083439616
I did look at the BlueMix platform, and it seemed like a big mix of various APIs I could tap into. Some of them looked neat, but I never had a reason to really use any of them.
How much is Bitcoin worth without translating it to another currency? Schrute Bucks and Stanley Nickels.
The person you're replying to is saying that all the hype around blockchain for contracts, identifiers etc. is most likely way overblown and we will look back in a few years and see it as a mostly useless fad
See why your argument doesn't prove Bitcoin is a good thing in the long term and/or if you're a believer in its superiority to centralized banking?
You could point it at your data and it would tell you the answer before you'd even thought of a question.
No wonder we're disappointed.
Is this standard practice in machine learning? This sounds more like regular programming to get exactly the outcome you want.
There are semi-supervised techniques to do stuff like this in a more systematic/automated way, but you still don't get anything for free: the outcome depends on the priors used to do the semi-supervised voodoo. In a generous moment I might assume this is what they meant, but it's still dumb.
The problem is that Watson was sold as a kind of "AI Doctor", which is pretty much ludicrous
Finally got my hands on it and was incredibly disappointed. Lacking features, archaic interface, and not delivering on what it was promised. Talked to some guys working for the Watson team and they told me that half of the AI stuff is just them doing it manually under the guise of "integration and configuration".
I did have fun "talking" to Watson whenever I was bored at work, trying to make him swear
I mean, GPT-2 still often produces nonsense more similar to dream imagery than useful reasoning, but I gather most medical residency students are half-asleep most of the time anyway, so... :)
Is there any example of gpt-2 in the wild that is not nonsense?
> Overpromised and Underdelivered
Sounds kind of catchy, to be honest.
It's super difficult to beat something like linear regression in this sort of thing (ideally combined with domain knowledge -something neural approaches mostly fail at), and even linear regression gives awful results.
The hard cases, which would be useful to a doctor, occur very rarely. My father's unusual reaction to a post-bypass drug regimen was something like the 3rd time that happened in Canada. How do you "train" that into a neural network?
To make this concrete: I used to do a lot of HIV testing and counseling. Whenever I'd ask someone "how often do you use condoms?" they'd 100% of the time say "every time!" When I'd then switch to asking "when was the last time you didn't use a condom?" they'd often reply along the lines of "last week."
This kind of issue happens not just with awkward questions, but also with more "objective" data like labs and medications.
- Was the antibiotic script written for a patient necessary for their condition, or was the physician tired and couldn't fend off a particularly assertive patient who was convinced antibiotics would solve their viral infection?
- Are labs randomly ordered, or ordered in a targeted fashion based on the "hunch" of a physician? If the latter, then the presence of lab result would itself be correlated with having some disease, and could thus throw off your entire model.
- How does the patient's insurance status influence the tests they have ordered? At one clinic I'm aware of, they typically order HIV/GC/Chlam screenings as a bundle, but if the patient has a PPO insurance they are also more likely to throw in a syphilis screening.
- How does patient preference factor into what gets diagnosed, ordered, and treated? You might have two early stage, equal risk breast cancer patients, one whom opts for a double mastectomy because she saw what happened to a coworker who had breast cancer and doesn't want to risk recurrences, whereas the other opts for a lumpectomy because she thinks her risks are low enough to be more conservative in treatment.
These issues are all solvable, but for AI to have a useful impact in healthcare, there's a lot more work than just throwing a bunch of data into a deep learning model.
Except for some really common problems, like cases of influenza, every patient is unique.
AI can be used to handle simpler problems - identifying patients at risk of repeat admissions, or flagging the likelihood of dehydration for example.