From Books to Knowledge Graphs
arxiv.org
arxiv.org
Only half joking. I do agree that AI content is going to change the landscape. However the things it does are all things it "learned" to do because... It was already present in the training data. The big change is just the scale those things can be done at.
with all these strides in AI, GPT, etc, does anyone know of papers about extracting into graph or other form of knowledge representations deeper semantics ?
for example instead extracting what the paper is about, how it relates to other papers, nodes and edges linking to field specific terms/concepts mentioned, translation of paper specific terminology to commonly agreed terminology, why it's citing said references (is it an example of use of a technique ? a proof for a claim ? something that's being refuted, etc ..), where in the text it's being used. and other kinds of "deeper meaning".
I've glanced at some papers that seemed to do something remotely close to this but it was restricted to doing NER for only medical/biology terms.
I think LLMs will be capable of orchestrating between systems and maintaining a conscious narrative but knowledge graphs will solve the update and references problem.
If interested, a few rabbit holes to explore (no affiliations):
https://scite.ai -> best option for citation mapping, but same issues you described above
https://www.semanticscholar.org and AI2 -> the best group working on tooling in this space
https://www.weave.bio -> early startup trying to build this out
The hardest challenge in my view is solving the intermediate representation issue. You have to establish a DSL/nomenclature that provides the range required to represent a complete scholastic discourse while also being computable.
Right, you'd basically be writing an interpreter for English
Generally, what I found useful to build a graph between "topics" or entities was to use a HyDE[1] prompt to generate possible distinct definitions and then build a nearest-neighbor network from that. This successfully identifies related concepts in the truly abstract sense rather than literal entities.
I haven't uses Wikidata info yet, but hoping to expand to that in 3 or 4 months.
HyDE sounds like an interesting approach. All dense retrieval approaches suffer from the problem you outlined in the blogpost. Have you looked at keyword-based or late-interaction models for retrieval such as ColBERTv2[1]? I find that late-interaction methods seem to offer best trade-off between semantic intelligence (precision) and retriving relevant documents (recall).
This semantic extraction can be extrapolated to most any trainable context. A useful one I've worked with involved mapping supplier-partner relationships. A well built supply chain graph can identify every layer of risk in a single supply chain and provide the provenance to back it up.
IMO, the biggest blocker to more mainstream use of Knowledge Graphs (even in the commercial world) is an actually intuitive interface for knowledge exploration. The real market innovation behind GPT isn't its 175 billion parameters, its the prompt interface that makes ChatGPT so universally accessible.
Also, if you want to scale, LLMs are going to prove to expensive, so eventually data need to be logged somewhere to create a fine-tuned model that can do sub-tasks and ideally do them better. What do you think?
Depends on what you mean by "better". With more accuracy?
"GraphGPT converts unstructured natural language into a knowledge graph. Pass in the synopsis of your favorite movie, a passage from a confusing Wikipedia page, or transcript from a video to generate a graph visualization of entities and their relationships.
Successive queries can update the existing state of the graph or create an entirely new structure. For example, updating the current state could involve injecting new information through nodes and edges or changing the color of certain nodes."
If anyone is interested in doing research properly and building a revolutionary public domain knowledge graph for your domain, I'm literally betting my house that this is the way to go.
TrueBase powers PLDB.com (140K lines of data), and our newer one CancerDB.com (13K LOD). A new one for physics and math is coming out this week!
https://try.scroll.pub/#scroll%0A%20comment%20This%20is%20wr...
Obviously a more useful example would be domain-specific knowledge in industry or medicine, but a generalized approach to ontological encoding from a given dataset would probably require a lot of interesting techniques and math that's way beyond my head.
But that's probably something easily done with current technology, what I'm interested in is learning how to talk about/learn about this concept. Distilling "knowledge," or "concepts" into "parameters", like, defining a DNA-like code for a given corpus of data... sorry I'm rambling but hopefully someone can relate
Example:
https://github.com/reidjs/simpsons-gpt/blob/master/output/ou...
The results were... not great. Maybe on the new GPT iteration it will be more successful though.
I am in a team developing the next generation of knowledge graph analytics, currently faster than almost all tools out there and its open-source.
The main barrier to entry, and issue, is that not many people know about them. Not many know how to use it, they dont understand the algorithms and thus are unaware of the benefits.
We should be careful to not treat the current hot thing as a hammer.
> Anytime you have an ML system where humans are designing how the information is organized (feature engineering, linking, graph building), it will work poorly.
This is a very broad assertion that I think is not accurate.
> The algorithm needs full control over the information at the lowest level (characters, words). Only then can something that truly works be built.
But data preparation and management is fundamentally important which means there's a human shaping the information at the lowest levels. There's no single ring to rule them all. This is also just fundamentally erasing the utility of any supervised model or hybrid system which is a bit silly.