IBM Sees Broader Role for Watson in Aiding Research
blogs.wsj.com
blogs.wsj.com
"Ford sees broader role for Taurus in aiding transportation networks of the future!" Of course they do.
I don't mean to be glib; and I do understand that "answering natural language questions" is distinct from "search engine". But I have the feeling that it's a capability distribution that is much more impressive in the context of a full-court IBM sales pitch to an executive who saw it on Jeopardy once and wants to get in on this machine learning thing that everyone's talking about, than as a tool used by experts to do a task or set of tasks better. Their shameless press release marketing isn't exactly dispelling that notion.
I can give more details if anybody is curious. (Disclaimer: I work for Watson)
A few questions:
1. It feels a lot like literature-based-discovery. Is it ? if so , how is watson better than current literature based discovery methods that did create some scientific hypothesis AFAIK?
If not please share more details(if you can of course) about how it reaches hypothesis).
2. Does it work well in non biological literature ? Are there any examples ?
It uses a lot of pretty advanced tech to do that: * State-of-the-art statistical annotators * Some deep domain knowledge understanding of the entities themselves (e.g., how chemical entities are composed) * Some nifty visualizations * Many component from the Jeopardy stack (you can read the Jeopardy papers for more insights) - deep question parsing, candidate generation techniques, scorers, etc. * Down the line it'll integrate more reasoning techniques along the lines of http://www.research.ibm.com/cognitive-computing/watson/watso...
2. Our first version was focused on life sciences and intelligence, we are expanding it to finance, law, and education.
More abstract: Watson, based on this article (http://www.wired.co.uk/news/archive/2013-02/11/ibm-watson-me...), can already diagnose cancer better than a human doctor. Will we see more of this functionality rolled out publicly? Or will it be reserved for use by healthcare organizations?
Needless to say I was a bit annoyed, because I was already using fact extraction in the system I wanted to test Watson's query 'skill' on. I'm in no position to store the raw text. That would require over a hundred times more storage, probably closer to a thousand times the storage costs, making it fiscally untenable for me to even build the database let alone a product with it. Any idea if release 2 in October will change this restriction and give people who aren't sitting on GB of unstructured text, like myself, a chance to apply.
The breakthrough with computational medicine is actually using it.
See https://www.youtube.com/watch?v=8lGJ0h_jAp8 for what the Watson medical diagnosis product looks like - it's not just about diagnosis, it's about evidences.
We even see it with the low acceptance of Watson in medicine and IBM's move into africa.
I wonder thought: why hasn't IBM offered a "second opinion" service directly to consumers ? It could might have sped adoption.
A "second opinion" service would have a lot of potential legal pitfalls.
I wonder, with the API , would it be possible for a startup to build such "second opinion service" , with reasonable costs ? will you block it?
Hint to IBM: at least make all documentation, example code, etc. available for perusal.
There have been many people working on similar projects to this in academia and otherwise. "Identifying the proteins that modify p53" is incredibly easy with the database information in Medline and doesn't require a "Watson" to do so. That being said, finding connections within medical literature could be interesting, but only if it is combined with other public databases that describe protein interactions and gene information. This blog post could be much more detailed in this sense.
As someone who works in the biological sciences and who has played around with various machine learning algorithms (ann, k-nn) this sounds like an advertising gimmick. I am all for using computers to process and analyze data, but frankly machine learning has not reached a point yet where algorithms can "identify" new hypotheses by sifting through the scientific literature.