> if you have a small graph would it unreasonable to use graphscope to analyze it - what if it was a small graph that you expected to grow big later?
You are right. GraphScope is built for processing very large graphs, but for small graphs, it should work just as well. If you expect your graphs to grow too big to be handled on a single machine in the future, I think it is a good idea to start with GraphScope even when it is small.
Stay tuned, we got plans to make GraphScope nicer to use for smaller graphs without a k8s cluster.
> What if you have a graph inside a graph db - I guess GraphScope gives you extra performance by allowing you to distribute processing for tasks that would make sense - any other benefits like that?
Besides extra performance, there could a few other benefits depending on your scenario.
1. Many graph analytical / GNN tasks takes a lot of computing resources. A single task can take hours to run even with 1000 CPU cores.
It makes sense to run such tasks in other machines/systems without adding too much burden to a graph db to avoid affect its quality of service.
2. Fully integration with Python makes it more flexible to do data analytics. For example, you can leverage the ability provided by numpy, pandas and mars (https://github.com/mars-project/mars) along GraphScope with zero-copy thanks to our storage engine vineyard (https://github.com/alibaba/libvineyard)
3. Besides distributed processing, extra performance can also come from the efficient graph layout in memory, and other optimizations on the compiler and runtime-level. GraphScope is ~100x faster on Gremlin, and even more on graph analytical algorithms like PageRank, compared with graph dbs like JanusGraph.