Google Knowledge Graph Search API
developers.google.com
developers.google.com
Google's knowledge graph is much, much better than any of the open data competition because they have done the work to make it consistent (not in the ACID sense but in the completeness sense)
For example, Wikidata appears good on the surface, but as soon as you try to build against it you find huge holes in the data.
As a more specific example, the most common example you will see on Wikidata is "list the cities with a female Mayor in order if population." Great, except it turns out that many (most?) cities aren't marked up with the attribute that makes them considered cities for the purpose of that query.
Knowledge APIs add typing to search. That's really important because it let's you disambiguate queries well (Apple computer vs Apple fruit) and behave more intelligently based on that type.
Things like the DDG API (mentioned in this thread) don't do that. DBpedia/Wikidata/Yago do it, but so inconsistently that the benefits are hard to make useful (as you are coding for the multiple ways types are handled).
The knowledge graph is much, much more than Freebase. The Freebase data is still available to download, and they are moving it to Wikidata.
http://www.cs.cmu.edu/~nlao/publication/2014.kdd.pdf (page 3, comparison table) KnowledgeGraph is from the same group as Freebase, just more data and closed.
A stale 6 months old Freebase dump gets more useless over time. Wikidata has a different license, it's just a PR piece that means little. 15,473,837 (Wikidata) vs 3,146,939,673 (Freebase) - little has changed since Jan 2015.
What's the issue with Wikidata's license? Both Freebase and Wikidata seem to be Creative Commons licensed. Is there a catch?
As for progress, if there's anything slowing the migration of Freebase data into Wikidata, I would guess it's Wikidata's different citation standards.
Given that the original Freebase data is still available, would it be correct to say the issue is Google isn't releasing their new data for free?
How about: Google shut-down a knowledge-base that was curated by a community, that provided regular data dumps, an online interface and an API - all with an open license (the original data source is nevertheless Wikipedia et al). Google's new venture is basically the same core technology and data but the crawler run also over the scrapped web content. And the only access for non-Googlers is via an API. Make your own conclusion from that.
Again, the Google Knowledge Base is much more than an expanded Freebase. It uses Google's Knowledge Vault project to extract from sources outside Freebase, as well as to evaluate and update the Freebase resources. To quote:
In particular, KV has 1.6B triples, of which 324M have a confident of 0.7 or higher, and 271M have a confidence of 0.9 or higher. This is about 38 times more than the largest previous comparable system (DeepDive [32]), which has 7M confident facts (Ce Zhang, personal communication). To create a knowledge base of such size, we extract facts from a large variety of sources of Web data, including free text, HTML DOM trees, HTML Web tables, and human annotations of Web pages. (Note that about 1/3 of the 271M confident triples were not previously in Freebase, so we are extracting new knowledge not contained in the prior.)[2]
[1] https://query.wikidata.org/#SELECT%20%28COUNT%28*%29%20AS%20...
{
"@type" : "EntitySearchResult",
"result" : {
"@id" : "kg:/m/0y49634",
"name" : "Software engineer",
"@type" : [ "Thing" ],
"description" : "Fictional Character"
}That's disappointing.
The second result in a query for "Rogan Josh" is "Jeremy Clarkson" a Person of "Top Gear" fame. Digging up the freebase record doesn't show any obvious reason why this would happen.
And stalled Freebase data from summer 2015 gets more useless every day.
https://en.wikipedia.org/wiki/Freebase
Paper: http://www.cs.cmu.edu/~nlao/publication/2014.kdd.pdf (comparision table on page 3, compare Google KnowledgeVault, KnowledgeGraph and Freebase)
Microsoft bought Powerset (for Bing and Cortana AI), IBM recently bought Blekko (for Watson AI). Google closed Freebase and reuse it for KnowledgeGraph (GoogleNow AI and Search). That recent development hurts independent AI research and smaller AI companies.
But seriously, I've been playing around with Wikidata's Query Service[0]. Here's an example...[1], the example asks, "What is `nature' a part of?" (Once you click through the URL shortener you can click execute to run the SPARQL[2] query. SPARQL's a W3C recommendation, sort of like SQL but for triplestores, but its details are not readily graspable I think.)
It seems like Google's Knowledge Graph is based on Wikidata? I used to think the Semantic Web was always going to be a decade away but now I think that it is going to play a large part in the near future of the web though if you pushed me to explain my change in reasoning I don't think I'd be able. What we need are Semantic Web Browser, no idea what they'd look like though :(
Here is Wikidata's table of properties from which it builds up its entire knowledge graph[3]. I think it's fascinating.
[0] https://query.wikidata.org/
[2] http://www.w3.org/TR/sparql11-query/
[3] https://www.wikidata.org/wiki/Wikidata:List_of_properties/Su...
I am very happy that Google opened up this API. I used to use Freebase, and I use DBPedia a lot. When I get home from traveling I am looking forward to kicking the tires of the new API.
Though, based on the amount of data Google has on the average user, and the fact that you have to sign-up to get an API key which is presumably associated with your search history, Gmail history (either any conversations sent from your Gmail account, or any mail you received dispatched from a Gmail account directed at you), they could easily determine if you meant Apple the fruit [you work for the USDA], Apple the company [you're an engineer in SF with a User-Agent history that's very heavily skewed towards Safari], or etymological basis of Apple, the word [you're a linguist], and disambiguate based on aggregate information. I'd imagine it'd be pretty trivial to do with their existing advertising profile + visit history of any site that either has Google Analytics or a Doubleclick ad.
[1] Again, I struggle to call it a graph, even if it's implemented as a GDB on Google's end, until the end-user traverses it, it's just a Knowledge API.
[2] https://developers.google.com/knowledge-graph/reference/rest... See: `types'.
Google definitely fuses user data into their knowledge graph. This is seen in Freebase's `g.` identifier [1]. I'm curious if they'd influence their publicly facing API algorithms using that data.
[1] https://groups.google.com/forum/#!topic/freebase-discuss/_8x...