Show HN: Graph Processing with Python and GraphBLAS
github.com
github.com
mm-ADT: A Multi-Model Abstract Data Type http://rredux.com/mm-adt/
Does this work in blocking or non-blocking mode? Naively I imagine there might be more opportunity for the GraphBLAS implementation to optimize execution in non-blocking mode.
Is there a way to efficiently store and load matrices to and from files? Ideally in a such a way that the data is just mmap'ed or copied directly into memory on load?
Does this only work with SuiteSparse or could it potentially work with a GPU implementation like https://github.com/gunrock/graphblast too?
Non blocking mode.
> Is there a way to efficiently store and load matrices to and from files? Ideally in a such a way that the data is just mmap'ed or copied directly into memory on load?
The Matrix.from_mm and Matrix.to_mm methods will read and write files in the standard Matrix Market format. They are wrappers around the underlying LAGraph functions.
> Does this only work with SuiteSparse or could it potentially work with a GPU implementation like https://github.com/gunrock/graphblast too?
It uses a lot of the GxB "extensions" that are to my knowledge only present in SuiteSparse. Investigating other implementations has been on my radar but not a high priority, certainly very open to getting PRs for helping with interchangeability!
[1] Sparse versus dense in GraphBLAS: sometimes dense is better http://aldenmath.com/sparse-verses-dense-in-graphblas-someti...
[2] RedisGraph: A Graph Database Module for Redis http://graphblas.org/?title=Graph_BLAS_Forum#Graph_analysis_...
[3] Graph Processing with Postgres and GraphBLAS https://news.ycombinator.com/item?id=19379800
NB: D4M was the original name before it was changed to GraphBLAS and became a standard.
[0] GraphBLAS: Building Blocks For High Performance Graph Analytics https://crd.lbl.gov/news-and-publications/news/2017/graphbla...
[1] A Billion Updates per Second Using 30,000 Hierarchical In-Memory D4M Databases https://arxiv.org/abs/1902.00846
Dask has a production grade distributed computing system (that is cloud compatible with kubernetes, yarn, EMR, Dataproc,etc).
Graph partitioning is a weird world, so will be interesting to see!
Search GraphBLAS https://hn.algolia.com/?query=GraphBLAS&sort=byPopularity&pr...
Does anybody here know about the advantages with respect to scipy.sparse ? Does scipy.sparse use graphblas internally?
graphblas provides a different bag of tools oriented toward solving graph problems, primarily around semirings and custom graph element operation. I don't think scipy.sparse, for example, provides a means for plugging in different semiring operations for computing shortest paths or max flow. On the other hand, graphblas does not provide decomposition or solving algorithms like scipy.sparse does.
It is my goal to make it so that scipy.sparse matrices, dense numpy matrices, and networkX graphs can be interchangeable used with pygraphblas in the most reasonably efficient way. This way we get all the tools in both bags. Would love to chat with more scipy developers to see where there can be unification.
UPDATE: it looks like there is some support for graph solving with scipy.sparse.csgraph, would be interesting to see how these compare to GraphBLAS based approaches!
[1] Tim Davis Research http://faculty.cse.tamu.edu/davis/research.html