KELM: Integrating Knowledge Graphs with Language Model Pre-Training Corpora
ai.googleblog.com
ai.googleblog.com
Wikimedia data is heavily depended on in the FAANG world for Google search, Siri, Alexa, etc... when Siri directly answers a factual question, I’d make a strong bet the answer ultimately comes from WikiDatas knowledge graph.
I just hope these companies give as much back to Wikimedia and society as the value they extract.
You could do this kind of graph -> text translation with conventional template-based tools, in fact people do that all the time. You very much run into the stages of "pick out a subgraph of salient facts", materializing text. If you scale it up you'll discover it has "erroneous zones" and end up building filters that block dangerous (likely to be wrong) outputs.
It's much harder to generate answers to questions. This calls for jointly choosing what knowledge to use in the answer and synthesizing text that presents that knowledge in a way that actually answers the question. This work is about this more dynamic problem.
Google does have an experimental API, but have not found an associated blog post or paper with it: https://cloud.google.com/ai-workshop/experiments/generating-...
If you already have a Knowledge Graph (KG) and want to populate its instances from documents, that's called KG Population, and Knowledge-net [4] is a good reference.
Relation Extraction is another interesting approach if you know which kind of relations you're interested in, OpenNRE [5] a good example.
[1] https://github.com/dair-iitd/OpenIE-standalone
[2] https://github.com/Lambda-3/Graphene
[3] https://github.com/uma-pi1/minie
If someone way more smart than myself could chip in on the subject, that would be pretty dang awesome.
In fact, the technique could probably be applied to language dialects as well (eg. Elizabethan, AAVE, etc.).
We talked about trying to do this at my last job.