I think this is how Sqrrl works or how they explained it when I was busy bombing interviews.
I think this is how Sqrrl works or how they explained it when I was busy bombing interviews.
As someone who actually built a graph based SIEM (Grapl) I have opinions here.
Relational databases are fine for graph work. So are non relational databases. What matters more is really how you can model your data and how easy it will be to work with in a "graphy" way, as well as how your system scales by default.
What a graphdb is going to do is get a lot of that "right" out of the box. You'll have a query language/ schema language that makes it easier to express graphs (annoying in SQL imo), that optimizes your edge indexing and join behaviors, that provides graph algorithms and optimizes for them (find a path, etc).
Typical relational databases are not going to be optimized for this work. You're going to need to be very careful about your indexes and queries to ensure that a join doesn't blow things up. You'll just have more work out of the box, and then it's just like... why didn't you use something that exists to solve this?
SIEM scale is much larger than people may expect. You have massive retention (years), billions of events a day, and lots of queries. This makes a lot of problems way harder - like intelligently indexing your edge data.
MATCH (a:Account)-[t:Transfer*]->(b:Account)
WHERE a.country = 'Canada' and b.name = 'John Doe II' and t.transfer_type = 'international'
RETURN a.ID
Which is… "show all accounts that have made international money transfers into John Doe II's account who lives in Canada. Source and target "table" is the same in this example.It is much simpler and more efficient to use a graph database especially to make a complex query where a leaf node circles back either to a "source" or to any intermediate node using the same or a new relationship. SQL will look nightmarish and will likely require manual hints to the query planner / optimiser.
Graph databases also allow to add (as well as delete) new relationships in an incremental and non-invasive fashion, as they become discovered – a useful feature for the analysis of a large or unknown dataset.
Graph databases require valid use cases, though, and they are not a generic substitute for a RDBMS (or for a random NoSQL database).
Steampipe [1] is an open source project [2] to live query all your cloud resources with SQL (via Postgres FDWs). It includes a dashboards as code layer written in HCL + SQL. We recently added support for relationship graphs & visualizations with hundreds of open source dashboards focused on cybersecurity [3].
In our experience, the hard part of these graphs was understanding which relationships / vectors are important to highlight and making them simple enough to browse in that way.
1 - https://steampipe.io 2 - https://github.com/turbot/steampipe 3 - https://steampipe.io/blog/release-0-18-0