Stay ahead of cyber threats with graph databases
memgraph.com
memgraph.com
What infosec has much more of is sales people. There's just a huge market trying to solve a tough problem. So from an outside perspective maybe you see a lot of bullshit but it's a highly technical field. Really, I only learned to be a software dev to up my infosec game, I find infosec to be a far more challenging and interesting field.
But maybe they too are salesmen types. Marketing their services to midwit management people on twitter. I include politicians in this too. Hurr durr china is going to defeat us in "cyber space" opens up a lot of business opportunities for government contractors. No one care about compilers, or operating systems the same way.
As for Twitter, that has to do with infosec having trouble communicating and a bunch of other stuff. I started to explain and it's not really worth getting into. Twitter was really important to infosec communication.
However, it’s also full of bullshit & marketing that want to convince you that every little technical detail is a vector for a potential cyberattack! You don’t want to be the next business suffering from ransomware gangs, do you? Better purchase our advanced endpoint detection powered by AI/“Neural Networks”, and don’t forget your automated security scanning designed with ChatGPT and quantum encrypted database backups to stop the Chinese/Russian state actors/hackers in their tracks!
In terms of the explosion in products, sure. But that doesn't apply to the technical people. There are lots of problems because there's a large market solving a complex problem. If you're actually talking to technical people they can be of a pretty high calibur.
Side note: you should read my other comment on this thread.
You'd want a lot of people with strong management skills.
I think this is how Sqrrl works or how they explained it when I was busy bombing interviews.
As someone who actually built a graph based SIEM (Grapl) I have opinions here.
Relational databases are fine for graph work. So are non relational databases. What matters more is really how you can model your data and how easy it will be to work with in a "graphy" way, as well as how your system scales by default.
What a graphdb is going to do is get a lot of that "right" out of the box. You'll have a query language/ schema language that makes it easier to express graphs (annoying in SQL imo), that optimizes your edge indexing and join behaviors, that provides graph algorithms and optimizes for them (find a path, etc).
Typical relational databases are not going to be optimized for this work. You're going to need to be very careful about your indexes and queries to ensure that a join doesn't blow things up. You'll just have more work out of the box, and then it's just like... why didn't you use something that exists to solve this?
SIEM scale is much larger than people may expect. You have massive retention (years), billions of events a day, and lots of queries. This makes a lot of problems way harder - like intelligently indexing your edge data.
Steampipe [1] is an open source project [2] to live query all your cloud resources with SQL (via Postgres FDWs). It includes a dashboards as code layer written in HCL + SQL. We recently added support for relationship graphs & visualizations with hundreds of open source dashboards focused on cybersecurity [3].
In our experience, the hard part of these graphs was understanding which relationships / vectors are important to highlight and making them simple enough to browse in that way.
1 - https://steampipe.io 2 - https://github.com/turbot/steampipe 3 - https://steampipe.io/blog/release-0-18-0
MATCH (a:Account)-[t:Transfer*]->(b:Account)
WHERE a.country = 'Canada' and b.name = 'John Doe II' and t.transfer_type = 'international'
RETURN a.ID
Which is… "show all accounts that have made international money transfers into John Doe II's account who lives in Canada. Source and target "table" is the same in this example.It is much simpler and more efficient to use a graph database especially to make a complex query where a leaf node circles back either to a "source" or to any intermediate node using the same or a new relationship. SQL will look nightmarish and will likely require manual hints to the query planner / optimiser.
Graph databases also allow to add (as well as delete) new relationships in an incremental and non-invasive fashion, as they become discovered – a useful feature for the analysis of a large or unknown dataset.
Graph databases require valid use cases, though, and they are not a generic substitute for a RDBMS (or for a random NoSQL database).
That said, neo4j+Bloodhound is something I am heavily invested in and not just me but most serious threat actors use Bloodhound+neo4j to choose lateral movement attack paths. Graph databases are how hospitals or govenrment departments get fully owned, except when they are 90s era flat IP network with everyone an admin over every server lol.
https://news.ycombinator.com/item?id=34342371 -- Bullshit Graph Database Performance Benchmarks
Yes, thank you for the reference. We're much appreciating comments, both positive and negative on the benchmarks, and are focused on even more workloads, and acknowledged datasets in the future
As for the title itself, I would not comment much on that.
The tldr; was that analysts are terrible. Powerful rule engines for the win. Also, accurately capturing and aggregating all the data in large organizations is very hard to do.
Our enterprise salespeople are on the way!