Show HN: Python library for embedding large graphs (Written in Rust)
github.com
github.com
It uses Maturin (https://github.com/PyO3/maturin) for this, which I've never heard of but sounds really useful.
The python parts are pythonic and the rust parts are "rusty". A joy to work with.
As an example: I created this to build this embedding of Mastodon instances. https://h4kor.github.io/fediverse-explorer/
As a next step I want to refine this as peering itself does not say anything about how much users of these instances interact. For this I will have to sample the users of instances and who they follow.
Also how well will this scale for large numbers of nodes, say thousands? Currently using networkx I have to perform tens of thousands of iterations on the spring_layout to get quite mediocre results.
I implemented this because networkx was to memory consuming and too slow.
Where networkx crashed with 24k nodes I was able to embed them with this library, it will still take several hours, but at least it's doable. The number of iterations should be on same magnitude as the number of nodes. For my graph #Iteration = 1/2 #nodes was sufficient.
The graphviz python library works great I've found with the sole exception of not having straightforward ways to edit the graph after creation. It's really more intended for packaging it up to pass to the command line application.
Will you have file import/export for .dot and similar?
I'm planning to better integrate it with other graph libraries, but at the moment you have manually translate between graph-force and you library of choice.
For speed I want to try two things in the future: Barnes–Hut simulation and doing the computation on the GPU.
I have spent some time looking into graph drawing algorithms and it seems to me that writing a good, optimised algorithm is non-trivial!
- Viz: Embedding gives x/y coords, and if they rendered the edges, a traditional graph view. A cool thing about recent shift to embedding approaches to the graph drawing problem is optimizing for objective functions that 1980's style spring layouts don't -- think the same things a neural network would optimize for. The code here appears more useful for small/medium graphs (ex: some ec2 tenant diagram) and I didn't see the neural network stuff, but with work, you can scale up several of the pipeline steps to handle 100X+ bigger ones (ex: we work with a lot of fraud or cyber event logs), and in headless cases, ~billion scale. It's a cool new subfield, google "graph drawing neural network".
- Decisions: Node, edge, and subgraph/graph embeddings are all super useful. I'm giving a talk at Friday's Infosec Jupyterthon (https://infosecjupyterthon.com/introduction.html) on the link prediction case for ~identity protection & resource ~monitoring (account takeover, insider threat, rogue devices, data leakage, ...) by mining log data, and as another example, link prediction is basically recsys, which is how any site with a shopping cart makes more money. Node classification, graph motif mining, etc are different but the same. Search for one of the many introductions to graph neural networks for a technical perspective, and I co-authored this survey at the beginning of the year to give a market perspective: https://gradientflow.com/what-is-graph-intelligence/
All this comes up a bunch in cyber/fraud/retail/supplychain -- we're certainly busy there. For anyone into that, we're hiring someone to own a bunch of backend/infra (k8s/gpu cloud/enterprise), 1-2 cleared folks in DC, and in Q1, (graph) data scientists. Simple question but one we're really into :)