What Is a Knowledge Graph?
neo4j.com
neo4j.com
While knowledge graphs are useful in many ways, personally I wouldn't use Neo4J to build a knowledge graph as it doesn't really play to any of their strengths.
Also, I would rather stab myself with a fork than try to use Cypher to query a concept graph when better standards-based options are available.
GraphQL is a JSON HTTP API schema (2015): https://en.wikipedia.org/wiki/GraphQL
GQL (2024): https://en.wikipedia.org/wiki/Graph_Query_Language
W3C RDF-star and SPARQL-star (2023 editors' draft): https://w3c.github.io/rdf-star/cg-spec/editors_draft.html
SPARQL/Update implementations: https://en.wikipedia.org/wiki/SPARUL#SPARQL/Update_implement...
/? graphql sparql [ cypher gremlin ] site:github.com inurl:awesome https://www.google.com/search?q=graphql+sparql++site%253Agit...
But then data validation everywhere; so for language-portable JSON-LD RDF validation there are many implementations of JSON Schema for fixed-shape JSON-LD messages, there's W3C SHACL Shapes and Constraints Language, and json-ld-schema is (JSON Schema + SHACL)
/? hnlog SHACL, inference, reasoning; https://news.ycombinator.com/item?id=38526588 https://westurner.github.io/hnlog/#comment-38526588
Not just that, w.r.t. reification they gloss over the fact that neo4j has the opposite problem. Unlike RDF it is unable to cleanly represent multiple values for the same property and requires reification or clunky lists to fix it.
> clunky lists
Not sure what the problem is here. The nodes and relationships are represented as JSON so it's fairly easy to work with them. They also come with a pretty extensive set of list functions[0] and operators[1].Neo4j's UNWIND makes it relatively straightforward to manipulate the lists as well[2].
I'm not super familiar with RDF triplestores, but what's nice about Neo4j is that it's easy enough to use as a generalized database so you can store your knowledge graph right alongside of your entities and use it as the primary/only database.
[0] https://neo4j.com/docs/cypher-manual/current/functions/list/
[1] https://neo4j.com/docs/cypher-manual/current/syntax/operator...
[2] https://neo4j.com/docs/cypher-manual/current/clauses/unwind/...
I don't think everybody should run away from property graphs for RDF or anything, in terms of the whole package they are probably the right technical call ninety-something percent of the time. I just find Neo4J's fairly consistent mischaracterization annoying and I have a soft spot for how amazingly flexible RDF is, especially with RDF-star.
In case you want to have a look a the SPARQL client I maintain, Datao.net, you can go to the website and drop me a mail. [i really need to update the video there as the tool has evolved a lot since that time]
that kid is 7 years old already, and in my understanding currently has only one active contributor. But idea of the project is very strong.
> While knowledge graphs are useful in many ways, personally I wouldn't use Neo4J to build a knowledge graph as it doesn't really play to any of their strengths.
I'd strongly disagree. The built-in Graph Data Science package has a lot of nice graph algos that are easy to reach for when you need things like community detection.The ability to "land and expand" efficiently (my term for how I think about KG's in Neo4j) is quite nice with Cypher. Retrieval performance with "land and expand" is, however, highly dependent on your initial processing to build the graph and how well you've teased out the relationships in the dataset.
> I would rather stab myself with a fork than try to use Cypher to query a concept graph when better standards-based options are available.
Cypher is a variant of the GQL standard that was born from Cypher itself and subsequently the working group of openCypher: https://opencypher.org/More info:
Do you by chance have any recommendations?
Which ones would you recommend?
The overall DX is quite nice. The apoc-extended set of plugins[0] make it very seamless to work with embeddings and and LLMs during local dev/testing. The Graph Data Science package comes preloaded with a series of community detection algorithms[1] like Louvain and Leiden.
Performance has been very, very good as long as your strategy to enter the graph is sound and you've structured your graph in such a way that you can meaningfully traverse the adjacent properties/nodes.
We've currently deployed the Community edition to AWS ECS Fargate using AWS Copilot + EFS as a persistent volume. There were some kinks with respect to the docs, but it works great otherwise.
It's worth a look for any teams that are trying to improve their RAG or are exploring GRAG in general. It's not a silver bullet; you still need to have some "insight" into how to process your input data source for the graph to do its magic. But the combination of the built-in graph algorithms and the ergonomics of Cypher make it possible to perform certain types of queries and "explorations" that would otherwise be either harder to optimize or more expensive in a relational store.
[0] https://neo4j.com/labs/apoc/5/ml/openai/
[1] https://neo4j.com/docs/graph-data-science/current/algorithms...
for how many records?
But here I mean "performance" in both retrieval time and the overall quality of the fragments retrieved for RAG compared to a `pgvector` only implementation. It is possible to "simulate" these types of graph traversals in pg as well, you'll have to work much harder to get the performance (we tried it first).
I'm always curious what people's use cases are with graph databases; do people find Cypher and SPARQL helpful? I've tried several times, but SQL is just so expressive. Postgres is still my favorite graph database (and CRUD RDBMS, and filesystem, and "data conversion tool").
What I have found is that "land and expand" using an index to find the landing spots is key for performance. Reason being once you "land" effectively, "expand" is cheap and fast.
Some of it will also come down to your graph design. If you have a lot of super dense nodes (analogous to a large JOIN), it will create a lot of memory pressure which it does not handle well.
But in a RAG use case, I don't see these as being issues.
Did you see the two blog posts that Tomaz Bratanic did on the topic?
For the ingestion: https://neo4j.com/developer-blog/global-graphrag-neo4j-langc... For the retrievers: https://neo4j.com/developer-blog/microsoft-graphrag-neo4j/
My general point on GraphRAG is that it extracts and compresses the horizontal topic-clustering across many documents and makes that available for retrieval.
And that by creating the semantic network of entities, you can use patterns in the graph structure to answer questions that rely on information coming together from different documents. Think the detectives board connecting facts with strings from many different sources.
Feel free to ping me for a deeper discussion: michael at neo4j
RDFox is a tool that uses Datalog internally. RelationalAI uses a datalog based approach. Another example is Mangle Datalog, my own humble open source project that can be found on GitHub.
The language in the article about relational being "non native graph" is a bit biased. With some developer attention, there are massive opportunities to store data in a distributed manner and with te right indices querying can be fast. Though to be fair, good performance will always need developer attention.
Take a bunch of tables and covert each row into a tuple (rowkey, columnName, value). Now take the union of all the tables.
^ knowledge graph
That’s it…but it’s not very useful yet. It becomes more useful if you apply a shared ontology during the import—ie translate all the columns into the same namespace. Suppose we had a “contacts” table with columns {“first name”, “last name”, …} and a “events” table with columns {“participant given name”, “participant family name”, …} — basically you need to unify the word you use to describe the concept “first name”/“given name”/whatever across all sources.
This can be cool/useful because you now only need one table (of triples) to describe all your structured data, but it’s also a pain because you may need to perform lots of self-joins or recursive queries to recover your data in order to do useful things with it. The final table has a very simple “meta” schema, and you erase the schema from each individual source so you can push the schema into the data.
When you try to grasp any complex topic your brain starts to build and connect a fuzzy network of topics and their respective positive or negative correlations and of course the weights between the connections.
Once you have unfuzzied the picture in your head you realize that the network is active and dynamic and that this network has different "modes" of operation and that some weights and correlations can change over time, while others are always static.
Mastering the dynamics of the knowledge graph is the final step in understanding it.
I've explained a similar thing to friends before, but I was always at a loss for relationships/insights that have led to concrete outcomes
E.g. telecom networks, electrical grids, travel networks (including logistics), etc.
All of these feature a very different node (usually a system or capital asset) vs an edge (usually a wire or constructed path), and insight into the structure is economically valuable.
(That's talking to more traditional uses of graph theory, less-so to modern knowledge graphs)
LLMs are great, but knowledge graphs are IMO indispensable to tame their shortcomings.
We had a similar problem, Datomic/Datascript not having an open format like RDF, but RDF being clunky and slow, so we build our own open-source solution in Rust (https://github.com/triblespace).
On an M1max we're currently at ~3us per query for a single result (so essentially per query overhead), and have something like 1m QRPS for queries with 3-4 joins.
I'm curious if you've somehow managed to shave off another order of magnitude, as I suspect that most WCO joins will be similarly limited by memory bandwidth. We for example worked out a novel join algorithm family (Atreides Join) and supporting trie based in-memory and succinct zero-copy on-disk data-structures, just to get rid of the query optimiser and its massive constant factor.
Prolog is better than datalog in a lot of ways: CLPZ, abduction, homoiconicity, being able to choosing search strategies for different problems, tabling, etc.
There's been some work to integrate prolog with LLMs:
https://swi-prolog.discourse.group/t/llm-swi-prolog-and-larg...
https://aneeshsathe.com/2024/05/10/dancing-on-the-shoulders-...
With LLMs enabling easy, if noisy, KG creation extracting knowledge into a computable form will lead to advances.
Drug discovery already uses the tech heavily, wouldn’t be surprised if it expands to more domains quickly now.
I have not yet found an application that combines all those functions and I have been considering to build one myself.
Knowledge Graph (disambiguation) https://en.wikipedia.org/wiki/Knowledge_Graph_(disambiguatio...
Knowledge graph: https://en.wikipedia.org/wiki/Knowledge_graph :
> In knowledge representation and reasoning, a knowledge graph is a knowledge base that uses a graph-structured data model or topology to represent and operate on data. Knowledge graphs are often used to store interlinked descriptions of entities – objects, events, situations or abstract concepts – while also encoding the free-form semantics or relationships underlying these entities. [1][2]
> Since the development of the Semantic Web, knowledge graphs have often been associated with linked open data projects, focusing on the connections between concepts and entities. [3][4] They are also historically associated with and used by search engines such as Google, Bing, Yext and Yahoo; knowledge-engines and question-answering services such as WolframAlpha, Apple's Siri, and Amazon Alexa; and social networks
Ideally, a Knowledge Graph - starting with maybe a "personal knowledge base" in a text document format that can be rendered to HTML with templates - can be linked with other data about things with correlate-able names; ideally you can JOIN a knowledge graph with other graphs as you can if the Node and Edge Relations with Schema and URIs make it possible to JOIN.
A knowledge graph is a collection of nodes and edges (or nodes and edge nodes) with schema so that it is query-able and JOIN-able with.
A Named Graph URI may be the graphid ?g of an RDF statement in a quadstore:
?g ?s ?p ?o // ?o_datatype ?o_langIt works great. I’m not a db expert but the flexibility and explicitness of the graph scheme clicks for me. It took me a while to come around on cypher but now that I’m there it makes sense.
A few years ago, a Good Samaritan on HN told me my bookmarks (I had about 10,000 at the time) were my "knowledge graph".
I had no idea what that was, but upon researching the concept, I was mind-blown by the simple truth of what I had been told.
Since then I became even more rapacious with my bookmarking (and especially editing their "Name" field to add tags and keywords), and I have about 30,000 bookmarks now.
And they truly are my knowledge graph. More so than the 1,000 or so text files where I store my notes on various topics. Mainly because of the Bookmarks Search feature. Let's say I want to refresh my understanding of Permutations, Combinations, Factorials (like I wanted to do, and did, yesterday). All I have to do is enter those keywords into Bookmarks Search and I instantly get the best (i.e. most relevant and intuitive for me) articles I've ever found on these topics (because I've bookmarked and tagged/keyworded them in the past). I find that more and more bookmarks end up with 404's these days, but then there's Internet Archive. Invaluable.
Which brings me to Chrome Bookmarks. There seems to have been no innovation in the last 15 years (other than the time they replaced the prior [and post/current] system with the horrible "Cards", and thankfully reversed course due to the howling protests and decided to instead offer "Cards" as a Chrome Extension [which is apparently not popular]). For example:
- One still cannot exclusively search folder names.
- One cannot search within a single folder only (by extension, one cannot search within selected multiple folders only).
- One cannot use Regex to search, or do any kind of Fuzzy Search.
- One cannot download (cache) a copy of a bookmarked page (in case the page 404s in the future) to store permanently in the some Zotero type storage system.
- When adding a new bookmark, one cannot conduct a search (using text keywords) for the folder one wants to store it in.
I realize that some of these features are available via third-party Chrome Extensions and the like. However, I stopped relying on third party anything after I relied on one for tagging bookmarks and this third party Extension decided one day to 404 on me (it shut down).
Please Google do more with Chrome Bookmarks.
Rant over.
(Addendum 1: Regarding the text files: yeah, I've tried Zettelkasten software. The one that came closest to my liking was FeatherWiki. However, in the end I continued with plain ol' text files because they're the most Lindy [plain text files are likely going to be the last format rendered unreadable by whatever destroys civilization as we know it].)
(Addendum 2: there seem to be some people out there who have uncanny storage and retrieval systems. E.g. @Balajis on X comes to mind. When he's in a debate/argument on X, he has an uncanny ability to pull out relevant contextual material, instantly, from his back pocket when the situation demands it. Let's hope they share their systems.)
I have now crosslinked the Bookmarks.
So while it may not be exactly what the linked OP article is talking about, it's a way of storing knowledge, linked via horizontal relationships.
But whatever floats your boat.
https://github.com/Fannon/search-bookmarks-history-and-tabs#...
If you're afraid that it goes 404: This extension is open-source, very easy to build and use locally and it does not make any external request or relies on external dependencies.