However, the AMD-GPU compatibility for CuPy is quite an attractive feature.
However, the AMD-GPU compatibility for CuPy is quite an attractive feature.
https://data-apis.org/array-api/
So it's possible to write array API code that consumes arrays from any of those libraries and delegate computation to them without having to explicitly import any of them in your source code.
The only limitation for now is that PyTorch (and to some lower extent cupy as well) array API compliance is still incomplete and in practice one needs to go through this compatibility layer (hopefully temporarily):
The last time I think this happen at market-scale was early 3d accelerator APIs? Glide/opengl/directx. Which has been a minute! (To a lesser extent CPU vectorization extensions)
Curious how much of Nvidia's successful strategy was driven by people who were there during that period.
Powerful first mover flywheel: build high performing hardware that allows you to define an API -> people write useful software that targets your API, because you have the highest performance -> GOTO 10 (because now more software is standardized on your API, so you can build even more performant hardware to optimize its operations)
https://scikit-learn.org/stable/modules/array_api.html
Disclosure: I'm a CuPy maintainer.
Also: You can just mix match all those functions and tensors thanks to the __cuda_array_interface__.
Code written for CuPy looks similar to numpy but very different from Jax.
As a sidenote, it is funny how this gets released in 2024, and not in say 2014...
Also CuPy was first released in 2015, this post is just a reminder for people that such things exist.
In 2024, with AI you can do these kind of projects very fast.
Is it changing though? Not only do PCIe interfaces keep doubling in performance, but CPU-GPU memory coherence is a thing.
I guess it depends on your target: 8x H100s across a PCIe bridge is going to have quite different costs vs an APU (which have gotten to be quite powerful, not even mentioning MI300a)
Last I checked (a couple months ago) it wasn't quite there, but I totally agree in principle. I've not gotten it to work on my Radeons yet.
I know you can enable ROCm for other hardware as well, but it's not supported and quite hit or miss. I've had limited success with running stuff against ROCm on unsupported cards, mainly having issues with memory management IIRC.
I think that every library required to build cupy is available in the universe repositories, though I've never tried building it myself.
[1]: https://salsa.debian.org/rocm-team/community/team-project/-/...
You could checkout some of EuroCC's courses. That should get you up to speed. https://www.eurocc-access.eu/services/training/