How to build a graph visualization engine
memgraph.com
memgraph.com
I think more people should build more graph visualization engines, but you're going to have a hard time competing with how slick pygraphistry is, but there is not enough alternatives that are worth a shit. Graphistry, when I left was able to do half a billion nodes (my memory might be fuzzy, could have been 250,000 but I think we ran tests and got 500k) all memory resident in the browser. Our goal was a billion. I'm not sure what it can do now but it has UMAP and some other very fancy features.
Here's a stupid unmaintained thing I made that got me the job at Graphistry. I did not build the graph vis engine, but I did write the procedural color stuff. It's ugly and terrible. Lol. I'm gonna release something some time soon that is a rework on that idea.
Also, I talk to Leo. If you have specific requests that you want to list, I'll make sure he reads this thread.
EDIT: Obviously when I say everything has GIS coordinates I'm not quite talking about outer space, but I think outer space is even covered by the rest of this idea. So is tiny space. This realization came to me when I was reading about people using PostGIS as a database for chemistry simulation. I have no link for that, I haven't thought about it in years, but now that I'm talking about it, I'll try to find it and post a link if I find it as I would like to readdress.
Scale - backend: As we're rapids.ai-native (helped start early days of both Apache Arrow + Nvidia RAPIDS.ai), we work with customers doing billion-level nodes/edges in interactive time on GPU servers. Mostly for fast ingest, ETL, + graph neural nets / manifold learning, and we're slowly pushing that into the visual stack.
Scale - frontend: We normally recommend reducing down to about ~2M edges or less. For sensible visual experiences, add in auto algorithms that cut to more like 500K. We've planned a way we think we can do another 100X, just not (yet) an engineering priority. Fun fact: your browser's JS VM is limited to ~1GB of RAM, so we're already at that limit in practice.
RE:Scaling graph visuals as an engineering practice, it looks like memgraph is starting where neo4j reached a few years ago, and makes sense. That approach doesn't really work well for the use cases memgraph advertises for, because as soon as a bunch of user/customer/IT/etc events happen and get visualized, the browser crashes. Optimization approaches like wasm and workers are clever -- a v0 prototype of graphistry did that! -- but we found is too unreliable to be the path for good performance across users of most operational teams ("Works on my machine" syndrome). Do it, but that shouldn't be the main source of 100X performance, just a 2X boost. We end up connecting GPUs in the browser to GPUs in the datacenter not just for 100X'ing this kind of stuff, but for a predictable performance way that limits how often your user's browser crash on real datasets.
Also maybe not obvious, this article focuses on interactive rendering, but a lot of the challenge after they figure out how to solve it is interactive analytics too. Most layout algorithms have non-linear complexity, so O(500K) is actually a challenge. A lot of our GPU offloading work nowadays isn't just rendering but layout, ETL, ML/AI clustering, etc. This ends up overwhelming the browser (why we do distributed GPU), and OLTP graph DB's aren't good at that either -- Neo4j basically had to write a V2 DB-in-a-DB to make their Graph Data Science module perform.
And no worries: There's no competition because Graphistry isn't a graph database :) Most of our users will do something like databricks dashboard / jupyter notebook / powerbi / etc query <> graphistry visual. I bet pairing a great streaming db like memgraph with Graphistry would combine respective engineering strengths quite well!
Did you mean a million? 500k is half a million not billion
Happy that people are building new graph visualisation engines and I think still a very open space for innovation in interactively exploring graph/network structures.
Cosmos is a really new lib, so when we did our initial research it wasn't even out yet. You can read more about our research over here: https://memgraph.com/blog/you-want-a-fast-easy-to-use-and-po...
From what we've understood, it's somewhat limited when it comes to graph styling. But an amazing technology nontheless.
Actually that was the main reason (along with the note that main authors are not contributing to visjs any more [1]) for a creation of the Orb where we fixed the blocking UI issue with graph simulation. Orb engine has two parts now:
* Simulator that doesn't depend on the DOM so we can move its heavy calculation to the web worker - we use d3-force for it [2]
* Renderer is pretty much influenced by vis-network, using similar style mechanism and canvas drawing capabilities (we credited vis-network in our code for those sections)
[1] https://github.com/almende/vis/issues/4259#issue-412107497
* first two sentences: You shouldn't build from scratch
* next four sentences: How not building from scratch failed for us
* entire rest of article: How we built from scratch and made an amazing product you should try.
I guess the implication is "you shouldn't build an engine because we just built the engine to end all engines"?
Even though "end all engines" might not be 100% correct, because the initial idea of the Orb is to make a single interface where the background (simulation and rendering engine) can be changed.
To add to your title, "you shouldn't build an engine because we just built the engine to unify all engines". :)
Business models that support network visualization: mostly, not such a great story. Customers want to solve problems, not just look at pictures of networks. Inevitably this drives the work toward domain-specific capabilities in areas like computer security, fraud detection or bioinformatics. It's a slippery slope. If you stay focused on core algorithms, your audience is other tool builders i.e. cost centers.
Scaling up network visualization: fascinating technical problem, but a human can't actually see a million objects at once or form a mental map of their locations. So it's more like a clustering problem. Not a big surprise that Graphistry adopted uMAP. It's treating nodes more like points in big plot. We're not concerned with the same problems as illustration quality rendering of small readable graphs.
Building your own: an appeal of network visualization is that you can get going by just writing some kind of physical simulation, assign reasonable coordinates to nodes, drawing edges as lines, and poof you're done. If your goal is consistently making concrete diagrams that look like a human drew them (with nodes that have shapes and ports, various kinds of labels, constraints on edge routing, nesting, aspect ratio control, etc.) there are so many intricate subproblems that you could spend years on any of them. But what's the financial incentive?
The research frontier: no doubt machine learning will eventually transform this domain the way it has many others. The combinatorial objectives of network diagramming make it challenging for now. (Can an algorithm learn orthogonal planar layouts with port constraints? Maybe. Would like to see that.) Another frontier is to extend general methods for declarative 2D layouts. People don't want just pictures of networks, they want more elaborate diagrams: computer networks, metabolic pathways, business processes, cryptocurrency transactions. Network visualization is only a subproblem in information visualization. This ties in to the first point, people need to solve problems in a specific domain.
I think ELK can do that: https://www.eclipse.org/elk/
When viewed through archive.org page takes >30 seconds to display any content other than menu's, orange back ground, and "are you having problems?" chat prompt.
Guess I'll have to work out how to build my own graph visualization engine :)
https://web.archive.org/web/20220916170642/https://memgraph....
No way in hell am I going to trust a company on any technical matter when they have to cobble their own product together through that many external services and trackers.
Except when you actually want to learn something. Worst possible advice to give someone and I haven't even read the second sentence of this blog post.
In the context of this article, I agree. Professional products should probably not try to build everything from scratch. But I do think it's important to acknowledge that building things from scratch as personal projects is one of the absolute best ways to gain a deep understanding of a topic. Even if you go on to use graph visualization engines built by others in all your future work, that doesn't mean you shouldn't give it a go on your own time if you're interested in trying it out.
Not trying to be contrarian or criticize the article, just my 2 cents. In general I agree with the sentiment of using existing solutions wherever possible.
Of course you won't write anything that you hate at first, because it's for fun. But hating something for some days or weeks is part of the fun, too. You are challenging yourself, not doing it the easy way, to appreciate how well the others are doing.
If you want to turn it into a business, so yeah, don't write a graph visualization engine, or crypto stuff, or anything really. Get a job, that can be fun too, and let the entrepreneurs figure out the rest.
I think you and I agree. I'm not being crass, but I'd if you had more thoughts on what I'm trying to say about what you're trying to say I'd be interested in further discussion.
It did a lot of performance wankery, like templating system (not my lib, I just used available one) being just Go code embedded in HTML that needed to be compiled with the rest of the app, or pre-generating HTML from Markdown on load. Using same language to write app as to write templates was nice tho.
Most of it was mostly "how far I can go without cutting features" because really going from 10ms per page to 0.2 ms per page has no difference after client RTT is involved.
Hell, even on localhost for some reasons Chrome always have few ms delay before starting download compared to any cli client, chrome shows anywhere between 3 and 8ms to start downloading, while FF sits at 0ms
I originally planned to have some fun with HTTP2.0 PUSH but, well, while I kinda believed on authority that it is useful for something it turned out to be entirely stupid idea which apparently nobody bothered to test before pushing it into standard so I didn't get to do it. Maybe I should try again with HTTP3.0 hints.
I'm asking this for rhetorical purposes. I know, I know deeply why. heh. I'm glad you came out of it ok. Thank you for your honest response.
I wanted something simple that was just markdown for actual content, no database or anything more fancy. I have considered static gen and even eventually migrated my old blog to it as archive (not in english, and cringe anyway). But it really started as "well, it looks like fun thing to do and I will be scratching the itch I had, why not".
Mind you, that was in 2012, the first version was in asynchronous(!) Perl, Ghost didn't even exist at that point (and I didn't wanted to touch JS anyway), let alone any other alternative. It even did respectable ~3ms per render of the page.
I rewrote it in Go to learn some stuff and in the process also yeeted comment processing and just farmed it out to externals (self hosted, not written by me) app, as that's probably the most annoying and thankless part of blog engine when you include all kinds of spam detection that would need to be written.
On the nearby cementery I also have unfinished Go z80 emu (I did learn a bit about using Dear ImGui from it) and it rewrite in Rust (because what's better first project than that?), just coz I wanted to see just how much faster C<->Rust is compared to C<->Go interop(answer = a lot, rendering part got from >4ms to below 0.5ms).
I did it to the point where it could run some code, and quite fast too, I think it was down to few ns per instruction, and 8 byte prog ran in like 11 ns, which I kinda didn't expect from Go. It didn't emulate instruction delay tho, which would be required to emulate it with peripherals.
Program decoder was just...256 byte array with function pointers generated out of operator list, which probably helped
Bait and switch blog posts (i.e that start potentially interesting and then half-way morph into a poorly disguised sales pitch) are so tiresome.
Just as a note, Orb is not there to compete with high volume graph visualizations like Cosmograph, Graphistry, Linkurious. It is more as a child from d3 and vis.js, which are great libraries, that uses d3 simulation and vis-like canvas rendering. We really liked what vis.js team did with the styling of the graph and how you can customize it - this is often a limitation for high volume graph visualizations.
We could also discuss about the analytics usability of seeing a graph with 1 billion nodes. It is definitely awesome, but it is too much data to grasp on as a user seeing it. Clustering or other graph algorithms would help. I think the question is: What is the maximum graph size (number of nodes/edges) when it becomes hard to get any useful visual information expect the graph global state? (e.g. seeing a bar chart with 365 columns (days) is harder to read than a bar chart with a smaller sampling, e.g. per week or month).
I don't know the answer to this, but maybe you will have due to your experience with graph visualizations.