MatX: Efficient C++17 GPU numerical computing library with Python-like syntax
github.com
github.com
Which should be compatible with any Kepler late architecture, as in 6xx models from 10years ago+?
Contribution guide is here: https://github.com/NVIDIA/MatX/blob/main/CONTRIBUTING.md
Like most benchmarks it really depends on what you want to do, and since it's a general library everyone might care about different things.
Versus the advert in the root Readme, which is impressive but gives no data on the pareto.
My system changes from nvidia-525 to nvidia-535 to nvidia-520 to nvidia-515 on a daily basis because I need to reinstall a different CUDA version just to try some new paper's code.
PyTorch did it right, it now ships with its own CUDA and doesn't take a shit about version what you have in /usr/local. Everything else should do the same.
Your 'cuda' packages should use '>=' not '=='
e.g. 'cuda-12' should depend on nvidia>=525 NOT nvidia==525
Depends: cuda-libraries-11-8 (>= 11.8.0), cuda-drivers (>= 520.61.05)
$ sudo apt show cuda-drivers
Package: cuda-drivers
Version: 520.61.05-1
Priority: optional
Section: multiverse/devel
Maintainer: cudatools <cudatools@nvidia.com>
Installed-Size: 7,168 B
Depends: cuda-drivers-520 (= 520.61.05-1)I wish nvidia would just release a `sudo apt install cuda-all` that just keeps you updated with ALL possible cuda versions simultaneously. I know it would be 20GB but that's fine, it's a drop in the bucket compared to the datasets I play with.
What you're describing with an "all" install can somewhat be accomplished with containers right now and none of the dependency problems.
Using a runfile is ditching the package manager on a package-managed system, instead of using the package manager correctly.
Is that even legal ?
Up to my understanding Cuda is covered by an EULA license that explicitly require end user agreement.
Bundling system dependency in python software is always a giant shit show. I would not name that "Everything should do the same"
Any other python library (randomly cuPy) that also ship its own Cuda, with its own different version, will segfault happily because you have duplicated symbols in a single process.
And:
(1) It is a solution only Nvidia can provide: Cuda is a binary. And they will certainly not do that just to please the python community.
(2) That just create an other set of problems just because python packaging sucks in the first place.
Lets just solve the initial problem, shall we ?
Are they even comparing apples to apples to claim that they see these improvements over NumPy?
> While the code complexity and length are roughly the same, the MatX version shows a 2100x over the Numpy version, and over 4x faster than the CuPy version on the same GPU.
NumPy doesn't use GPU by default unless you use something like Jax [1] to compile NumPy code to run on GPUs. I think more honest comparison will mainly compare MatX running on same CPU like NumPy and for GPU comparison focus to compare vs CuPy.
I addressed it more here: https://news.ycombinator.com/item?id=37760120
So the test seems to compare one thread of a CPU launched in 2016 running a port of 40year old fortran code to a optimised FFT library on a GPU launched in 2020.
On newer GPUs, though, we have this huge L2 cache which makes the calculus a little different if your working set fits into it. e.g. Ampere A100 has 40MB L2$.
Is the UCLA/Nvidia/Raytheon collaboration (as presented in a recent GTC talk) a major force behind the development of MatX?
And then maybe also a comparison to Flashlight (https://github.com/flashlight/flashlight), xtensor (https://github.com/xtensor-stack/xtensor) or other C/C++ based ML/computing libraries?
Also, there is no mention of it, so I suppose this does not support automatic differentiation?
We don't support automatic differentiation (yet).
You might as well say they're free to make monitors that don't rely on color.
Although I'm not expecting on that scale here, but what's the planned future work here? E.g. matrix free operators and operations, sparse matrix support, or symmetry aware eigenvalue / singular value decompositions? There are official cuda libraries that support these (e.g. cuSPARSE).
Otherwise this looks good! Documentation is decent too.
Since MatX was originally intended for streaming/real-time processing the focus has been on making C++ for CUDA easier to use for those kinds of applications.
There are also a lot of things in cuPy and sciPy that don't make a lot of sense to do on the GPU, like offline tasks such as filter design in signal processing.
Also, since C++ users typically are writing in that language for maximum performance, we've put a large focus on making sure we are as close to writing optimized CUDA as possible. In general, most workloads we've tested are about 3-4x faster than the cuPy counterpart due to better fusion and language overhead.
We have discussed supporting cusparse as well, but cusparse typically requires a different input type like a CSR matrix. This is not something you'd typically want to detect/convert, so we're still discussing ways of getting this integrated cleanly.
We do have several versions of SVD, including one that calls cuSolver.
Feedback is always appreciated!
For those of us in-kernel day-laborers I can't wait for cuBLASDx and cuSolverDx releases, and for my brain to wrap around cutlass and cute. Please don't forget us!
I feel this effort is poorly differentiated compared to previous efforts.
That being said, in the example on the home page there are notable differences:
1) we have the run() method. The reason is that the expression before the run is lazily evaluated for performance and does not execute anything. Having the run() method allows you to run the same line of code on either a CPU or GPU by changing the argument to run()
2) in MatX memory allocation is explicit. Python does it as-needed, but this causes a performance penalty with allocations and deallocations that are not under your control. Specifically in the FFT example, numPy will allocate an ndarray prior to calling it, but on the same line. In MatX the allocation is (typically) done before the operation so you can control the performance of the hot path of code.
If you have any specific suggestions, we would love to hear it
MatX has been used in several projects with hundreds of microsecond deadlines, which is not usually something you'd choose Python for out of the box.
We are instead targeting users who already have python or high-level code that they need to port to c++ for whatever reason, and want to do it in the easiest way possible. With c++17 we're able to provide a simple syntax without compromising on performance compared to most native code.
see this radar pipeline example for one:
https://github.com/nvidia-holoscan/holohub/tree/main/applica...
The syntax of MatX also allows us to do more kernel fusion in the future to improve performance without any changes.