35 karma · joined October 3, 2023
Depends: cuda-libraries-11-8 (>= 11.8.0), cuda-drivers (>= 520.61.05)
We don't support automatic differentiation (yet).
Like most benchmarks it really depends on what you want to do, and since it's a general library everyone might care about different things.
see this radar pipeline example for one:
https://github.com/nvidia-holoscan/holohub/tree/main/applica...
Contribution guide is here: https://github.com/NVIDIA/MatX/blob/main/CONTRIBUTING.md
What you're describing with an "all" install can somewhat be accomplished with containers right now and none of the dependency problems.
I addressed it more here: https://news.ycombinator.com/item?id=37760120
MatX has been used in several projects with hundreds of microsecond deadlines, which is not usually something you'd choose Python for out of the box.
We are instead targeting users who already have python or high-level code that they need to port to c++ for whatever reason, and want to do it in the easiest way possible. With c++17 we're able to provide a simple syntax without compromising on performance compared to most native code.
The syntax of MatX also allows us to do more kernel fusion in the future to improve performance without any changes.
That being said, in the example on the home page there are notable differences:
1) we have the run() method. The reason is that the expression before the run is lazily evaluated for performance and does not execute anything. Having the run() method allows you to run the same line of code on either a CPU or GPU by changing the argument to run()
2) in MatX memory allocation is explicit. Python does it as-needed, but this causes a performance penalty with allocations and deallocations that are not under your control. Specifically in the FFT example, numPy will allocate an ndarray prior to calling it, but on the same line. In MatX the allocation is (typically) done before the operation so you can control the performance of the hot path of code.
If you have any specific suggestions, we would love to hear it
Since MatX was originally intended for streaming/real-time processing the focus has been on making C++ for CUDA easier to use for those kinds of applications.
There are also a lot of things in cuPy and sciPy that don't make a lot of sense to do on the GPU, like offline tasks such as filter design in signal processing.
Also, since C++ users typically are writing in that language for maximum performance, we've put a large focus on making sure we are as close to writing optimized CUDA as possible. In general, most workloads we've tested are about 3-4x faster than the cuPy counterpart due to better fusion and language overhead.
We have discussed supporting cusparse as well, but cusparse typically requires a different input type like a CSR matrix. This is not something you'd typically want to detect/convert, so we're still discussing ways of getting this integrated cleanly.
We do have several versions of SVD, including one that calls cuSolver.
Feedback is always appreciated!