DeepDive – System that enables developers to analyze data on a deeper level
deepdive.stanford.edu
deepdive.stanford.edu
So many open source "Knowledge Graph"-y type projects concentrate on building them like databases, with a query language that assumes the data in them is correct. You see this in things like Freebase, DBPedia and Wikidata, where they typically end up in a triple store and you query using SPARQL.
This isn't how the real world works, and there isn't a lot of publicly available software that takes this into account. There aren't even than many papers about it (the Microsoft Probase paper is one, and there is work from Florida University(?) about using Markov chains to reason while taking probabilities into about).
I'm excited to take a look at this.
I hadn't heard about DeepDive until now, but I did previously come across another project that does probabilistic reasoning: http://alchemy.cs.washington.edu. I cannot speak to how they compare since I haven't looked into either in great depth yet.
There is a lot of work on probabilistic inference. Huge volumes.
But then if you are interested on papers on using probabilistic inference on unstructured information there is a lot less information.
If you are looking for actual implementation reports there isn't much. The most interesting are the previously mentioned Probase paper, The Google Knowledge Graph paper and assorted IBM Watson stuff.
I've just realised that DeepDive is also the Wisci project[1], which is something I've been watching for a while
That Alchemy implementation requires a structured database of information, and doesn't really deal with how to get that.
These were some of the major projects: https://homes.cs.washington.edu/~suciu/project-mystiq.html, http://maybms.sourceforge.net/, http://infolab.stanford.edu/trio/, http://www.cs.umd.edu/~amol/PrDB/, http://dl.acm.org/citation.cfm?id=1376686.
I thought BlinkDB was data warehousing on hdfs? I dont see any mention of inference-like features in the docs?
Although some terms come up in both places (e.g., confidence bounds, noise, etc), BlinkDB and probabilistic databases are fundamentally different from each other (I have worked on both topics).
You're talking about the work that Prof. Daisy Zhe Wang and her students are doing over at the DSR lab. Go gators!
I think the answer is probably closer to weeks to months if working with field experts, depending on how deep you want to go.
The core of it is open source.
I think the most exciting thing about it is it brings more sophisticated computation to the more qualitative sciences.
The key issue in estimating how big a job it is is how complex your entity extraction and inference rulesets are.
What's included in the non-core parts? Are there patents to be licensed?
Firstly: IBM is increasingly using the Watson brand for things that don't appear directly related to the Jeopardy winning system (eg, Watson Analytics). When I talk about Watson here I mean the Question Answering (QA) system.
At a very high level DeepDive consists of a Knowledge Graph construction tool, and a probabilistic querying tool. Compared to Watson it is missing a natural language question parsing tool, and any way of dealing with questions that aren't in the KG.
Watson has (very strong) natural language understanding for multi-claused questions, and the Jeopardy version can do things like understand puns. Deepdive doesn't have anything comparable. In the open source space, the closest thing I'm aware of is SEMPRE[1][2].
Watson also has a evidence scoring module, and my understanding is that this can work against unstructured data. Deepdive doesn't have this, and instead relies on probabilistic inference. This is an excellent approach, but relies on doing content extraction first (ie, extract entities and relationships from text and/or other sources). The Microsoft Probase[3] group has published lots in this area.
[1] http://www-nlp.stanford.edu/joberant/homepage_files/publicat...