Ex: Preprocess many cols with scikit sparse (genomes, many-featured fraud, ...) -> cudf (= gpu dataframes) for most of workload -> infer relationships with gpu UMAP then gpu k-nn -> graphistry for visual analysis.
We're starting to experiment with bigger-than-memory (TB etc) scale via the gpu-enabled dask libs. I definitely recommend trying. If interested in adding visual analytics here, these'll be hitting our free + self-hosted layers soon, and drop a contact method if you'd like early access. Exciting times!
The main implementation is SuiteSparse::GraphBLAS, a C library which has two Python bindings (search for grblas or pygraphblas). Disclosure: I'm the author of grblas. Both are available on conda-forge for easy installation.
If you want to try it out without installing, here is a binder link: https://mybinder.org/v2/gh/metagraph-dev/grblas/HEAD?filepat...
What's the underlying data structure of the graph?