Maybe I'm just not skilled enough to handle such a workload with Postgres. But Neo4j handled it easily
Graph databases have a very narrow usecase, and it's almost always in relation to people - at least ime.
Though the data type isn't really important for the performance question, the amount of data selected is. So a 6-level depth graph of connections that only connect 2-3 entities would never get into performance issues. You'd be able to go way beyond that too with such a narrow window. (3 entity Connections on 6 level join would come out to I believe ~750 rows)
If you're modeling something like movies instead, with >10 actors per production you're looking at millions of rows.
https://github.com/joelonsql/graph-query-benchmarks
I haven't tried the latest versions of both databases though.
Any sense if ClickHouse can scale to several TB of data, and serve giant queries on the last hour’s data across all sensors without getting OOMKilled? We’re also looking at some hacked together DuckDB abomination, or pg_duck.
The nice thing is, once you understand the data model it becomes very easy to predict if it will fit your use case or not as there is really no magic to it.
What does your testing strategy look like with bigquery? We use snowflake, but the only way to iterate and develop is using snowflake itself, which is so painful as to impact the set of features that we have the stomach to build.
My current headache is what to do with an actually big table, 25 billion rows of json, for development. It’s going to be some DBT hacks I think.
God help you if you want to unit test application code that relies on bigquery. I’m sure there are ways but I doubt they don’t hurt a lot.
What are you doing with that JSON? What’s the reason why you can’t get a representative sample of rows onto your dev machine and hack on that?
I presume that for larger real world graph datasets, maybe there's some better algorithms and storage methods. I couldn't figure out neo4j fast enough, and it wasn't clear that I could map all of the stuff like block modeling into it anyways, but it would be very useful for someone to figure out a better production ready storage backend for networkx at least where some of the data could be cached in SQLite3.
This reminds me of "Just Use Postgres for Everything": https://news.ycombinator.com/item?id=33934139
My first thought is sharding and materialized views.
For timeseries I'd argue with timescale there is very few use cases that ever outgrow postgres. But I might be biased.
If you plan for it, for example store your graph in an export friendly format it should be ok.
Using postgres as long as you can get away with it is a no brainer. It’s had so many smart people working on it for so long it really comes thru. It’s nice to see the beginner or intro level tools for it evolving quickly similar to what has made MySQL approachable for so many years.