Diffbot launches AI-powered knowledge graph of 1T facts
venturebeat.com
venturebeat.com
This is such an over ambitious problem...
1. Webpages are ladden with errors, how do you deal with this? 2. Knowledge does not fit in a graph. It's asymptically a graph, as in: I can define relationships like: This recipe contains carrots. carrots contain sugar => this recipe contains sugar. Cool. But, what about this "sugar-free carrot cake recipe?" Well it still contains carrots, so still contains sugar... Contradiction? => requires human curating... 3. It doesn't even solve a real problem... Look at IBM watson, it probably knows a lot more crap than diffbot, and yet, is a pretty useless piece of software...
As you point out, the concept "sugar-free" does not equate to there being no "sugar" in the dish. Perhaps it means to folks something closer to that fact that no added sugar was explicit in the recipe. The word "sugar" itself has many meanings (the well-known problem of word Polysemy). So there are several similar but distinct concepts at involved here and an entity could have the value of true for one of these and false for another.
And while it may ultimately be doomed to partial failure, it´s still a beautiful thing to watch.
- OpenCyc http://www.cyc.com/opencyc/ - DBPedia https://wiki.dbpedia.org/
The Diffbot Knowledge Graph is built by applying computer vision and natural language processing techniques to reading all the pages on the web (which can be in any structure and human language) and extracting it into a structured form, without the element of human annotation in the build pipeline.
Obvious one is scale, Wikipedia has on order 10M entities and represents the work of thousands of humans whereas the Diffbot KG has 10B entities and is discovering about 120M each day, and is largely limited by the number of machines running the algorithms in the datacenter. The properties and facts indexed about each entity are also a superset because it is not limited to those that would be worthwhile for a human to curate. Lastly, it can be more accurate than facts found in a single source because the automated system utilizes multiple sources of that fact found across the web to estimate a probability of the accuracy of the fact.
The result is that you have a Knowledge Graph that is more useful for work and business because they are the entities you interact with day to day, not the "head" entities that optimize for popularity and the constraints of human curation.
Further, for this to scale, Diffbot has to have a way to align their entity IDs with IDs from other notable graphs like Wikipedia, Wikidata, Freebase, Wordnet or even Yelp, and the like, otherwise the data could be potential of diminished value.
How would I know that the "Cardi B" that's in my database with ID 321 and wikidata ID Q29033668 is the same as Diffbot's "Cardi B" with ID 561?
If we're talking notable people/things, you'd get more structured meaning out of searching a local Wikidata graph.
fantastic work here! As someone who's really excited about a machine readable web and have been working on it, this is fantastic! Unfortunately, while the Semantic web was to tackle this, the real life proliferation of the Semweb has been, atleast to me personally, extremely disappointing.
So this is a fantastic initiative, personally for me to know about.
Is there a plan to expose this data via a dev API of some sort for enthusiasts like us?
Say a SPARQL or even (Open)Graph API perhaps?
My experiences consulting with and working with companies interested in the domain has been that monetizing this data is extremely hard both legally and quality wise.
"Nike Tanjun near me" is a query fraught with danger. People typing this query want to find a retailer in their vicinity that sells this Nike product, but where do we source that inventory list from and how do we get our cut?
Before people start talking about DSPs and SSPs, this is a very different problem at hand.
To know that Nike Tanjun is a shoe sold by Nike, an ontology needs to exist that captures this knowledge so that the user's query can be decoded.
How will that ontology be sourced? Further, for it to be usable commercially, Nike has to agree to that encoding. Therin begin the challenges. If we encoded Nike Downshifter to be Tanjun, by mistake, then the user bought them based off our results, disliked them expecting the Downshifters to be like Tanjuns, we have an issue and Nike could persue the matter because we mislead the customer and affected their branding.
My primary clients are search companies or companies that want to provide rich search functionality: Google, Bing or even DDG do a phenomenal job in this space and the barrier to entry is pretty high.
So knowledge quality, mainteanaanace, versioning and temporal resolution ("The President of the U.S.", "The iPhone" are different entities over time) aside, is diffbot going to monetize this knowledge only as a B2B offering/addon to their clients or are there other "bigger" plans to monetize this tremendous undertaking and keep it rolling in the future?
AllegroGraph is a well known solution to this problem but out of reach of most enthusiasts like me, so if DiffBot came up with their own solution, I, among others would love to read about it more