An Introduction to Knowledge Graphs
ai.stanford.edu
ai.stanford.edu
Also remember that Wikidata is open source and you can fire up your own knowledge graph as docker containers on your laptop: https://wikiba.se/
If you have been disappointed by RDF-based technologies before, I would say Wikidata/Wikibase have significantly innovated on top of them. For example, they allow each statement to have qualifiers, references, depreciation/preferredness attached to them in a user-friendly way while also keeping simple queries simple.
For example, asking the same person the same question may yield different answers based on their mood or other environmental or situational factors. Who's asking the question can also matter, as does the specific phrasing of the question.
The reason for that is twofold:
1. Many of tools created for reasoning are research-first tools. Some papers were published about the tool and it really was a petter and more scalable tool than anything before it. But every PhD student graduates and needs to find a job or move to the next hyped research area 2. Tools are designed under the assumption that the whole ontology, all the instance data and all results fit in main memory (RAM). This assumption is de-facto necessary for more powerful entailment regimes of OWL.
Reason 2 as a secondary sub-reason that OWL ontologies use URIs (actually IRIs), which are really inneficient identifiers compared to 32/64-bit integers. HDT is a format that fixes this inneficiency for RDF (and thus is applicable to ontologies) but since it came about nearly all reasoners where already abandoned as per reason #1 above.
Newer reasoners that actually scale quite a bit are RDFox [1] and VLog [2]. They use compact representations and try to be nice with the CPU cache and pipeline. However, they are limited to a single shared memory (even if NUMA).
There is a lot of mostly academic distributed reasoners designed to scale horizontally instead of vertically. These systems technically scale, but vertically scaling the centralized aforementioned systems will be more efficient. The intrinsic problem with distributing is that (i) it is hard to partition the input aiming at a fair distribution of work and (ii) inferred facts derived at one node often are evidence that multiple other nodes need to known.
loose from modern single-node However, the problem of computing all inferred edges from a knowledge graph involves a great deal of communication, since one inference found by one node is evidence required by another processing node.
[1]: https://www.oxfordsemantic.tech/product [2]: https://github.com/karmaresearch/vlog/
Perspectivism is the understanding that it's impossible to interpret semantic "knowledge" without knowing the limitations and implicit context carried with the fallible, partial transcription of the truth (set in a world that obeys quantum mechanics, for one thing) into words (which run to about 1MB tops before the author gets bored.)
Quantum mechanics actually says there's less information needed to describe X area of space than classical physics implies there is.
[1]: https://link.springer.com/article/10.1007/s11406-021-00371-1
See also his talk https://www.youtube.com/watch?v=ujMgQqp8YSY
Stuff like schemas and data dictionaries and reuse are chinese finger traps for us geeks. Exquisite problems we can't look away from.
I eventually decided to treat most data ingestion (ETL) as screen scraping. Honoring Postel's Law. Pull out the interesting relevant bits as needed. Ignore the rest.
There's still an internal model, natch. But it's the smallest, most obvious model to support my immediate use cases. Nothing more.
Any NLP application directly encoding knowledge (objects, perspectives, whatever) is going to have major scaling constraints, since it's impossible right now to encode human-level understanding automatically.
what does this mean in this context?
Two columns is for the reader's benefit - your eyes can keep their place on the page much more easily when jumping half the distance to the beginning of the next line
We blew it, making a global effort to perfect fixed format technical typesetting rather than flowable text. Technical flowable text is just now becoming viable, and it is certainly not the norm for journal articles.
I've worked with systems where the relationships are typed and can have attributes just like the vertices allowing the system to model data in a more intuitive fashion.
You could then choose to add labels as either an additional function label: E -> L, or as several graphs layered on top of one another (especially helpful when you view each graph as an arrow N -> N, which in turn makes more sense when you have several different classes of nodes).
If anyone's interested I'm loosely basing this description on the description of petri nets as outlined in Tai-Danae Bradley's interesting (and readable) Applied Category Theory: https://arxiv.org/pdf/1809.05923.pdf
I've created some useful output from Wikidata's dataset with some-space saving decisions like focusing on certain languages and an arbitrary data structure. The dump is quite big and space at a premium.
But I don't know if it scales up to Wikidata size.
Yes because the web is already a kind of knowledge graph but it's mostly written in natural language, and thus it's very hard for machines to traverse and reason about it. The Semantic Web was an attempt to formalize some ways to make the web's inherent knowledge graph nature more explicit and thus easier for programs to understand.
No, because knowledge graphs predated the web by many years and KGs are a bigger topic than just the web.
Programming languages are valued both when they reveal inevitable design, and when they enjoy widespread adoption. We continue to have many programming languages, because there is no consensus on inevitable design.
The design of "this" reveals no deep secrets about the nature of the universe; the only parts that seem inevitable are the parts that seem obvious. And all of what one sees seems obvious; the choices involve what one doesn't see, what is left out of the system.
The value here is in widespread adoption. The system is good enough that people can agree to use it.
https://github.com/simongray/clojure-graph-resources#datalog
In principle, all of this information is available through the SPARQL endpoint or as an RDF export (there is also the simplified export that contains only “simple” statements lacking all of that metadata), so reasoning over this data is not entirely out of reach, but the sheer size (the full RDF dump is a few hundred GBs) is also not particularly practical to deal with.
https://en.wikipedia.org/wiki/Wikidata
Thanks for that! TIL.
It seems a fascinating project in epistemiology!
If you mean autocomplete UI or tooltips, look no further than the query editor and its Ctrl+Space at https://query.wikidata.org/
*) Here is an example of the number of attributes for a single entity: https://www.wikidata.org/w/api.php?action=wbgetentities&ids=....
Nodes = Tables
Edges = Foreign keys
Edge labels = Foreign key constraint names
it does fall down for graphy tasks like multihop joins, connect the dots, and supernodes. So for GB/TBs of that, either you should do those outside the DB, or with an optimized DB. Likewise, not explicitly discussed in the article, modern knowledge graphs are often really about embedding vectors, not entity UUIDs, and few/no databases straddle relational queries, graph queries, and vector queries
These can always be accomplished via recursive SQL queries. Of course any given implementation might be unoptimized for such tasks. But in practice, this kind of network analytics tends to be quite rare anyway.
One should note that even inference tasks, that are often thought of as exclusive to the "semantic" or "knowledge" based paradigm, can be expressed very simply via SQL VIEW's. Of course this kind of inference often turns out to be infeasible in practice, or to introduce unwanted noise in the 'inferred' data, but this has nothing to do with SQL per se and is just as true of the "knowledge base" or "semantic" approach.
but it think "good fit" is a stretch. when designing systems you generally want to look at data access patterns, and pick a data exec approach that aligns to that.
in tech, unfortunately, RDBMS are the "hammer" in "if your only tool is a hammer then every problem looks like a nail."
The one thing that's definite is that SQL is a bad choice for particular kinds of queries, though most graph databases don't seem to go much further than improving (?) the syntax a little bit and adding transitive closure (which is also present in several SQL databases). A few graph databases do allow for more complex (even arbitrary) inference, but this somehow never seems to make the headlines.
How so? What's practically missing?
Basic (non-recursive) SQL: Finite paths
Recursive SQL / Transitive closure (with negation): Regular paths
However the hierarchy of languages doesn't end there. And some graph databases allow you to add arbitrary grammar rules, which make it possible to add some complex rules like:
- If condition X,Y and Z is satisfied then person A and B are the same person
- Equality is transitive
- If two people are equal then each property of one is also a property of the other.
The part that makes this tricky is that figuring out two people are equal can cause other people to now suddenly satisfy the condition for equality.
You could also construct some examples by creating conditions which aren't "regular" (e.g. person A and B have an ancestor which is the same number of generations back for both of them).