CuPy – NumPy-compatible matrix library accelerated by CUDA
cupy.chainer.org
cupy.chainer.org
Along with a colleague I used CuPy in first Chainer and then PyTorch for implementing the Quasi-Recurrent Neural Network (QRNN) which at the time was far faster than even NVIDIA's optimized cuDNN LSTM whilst getting the same (or better) performance for many tasks.
CuPy at the time was both the easiest and most Pythonic of potential solutions for that problem - even if it did involve writing CUDA in Python strings =]
n.b. Our use case was literally pushing state of the art in research - CuPy is even more Pythonic if you're hitting more standard use cases.
I continue to have more sympathies for chainer.
Th autograd library was inspired by Chainer's design and took a lot of concepts (but not code) directly from Chainer. The neural network API is a bit of a hybrid. It's built on top of the autograd library but the layer names, implementations, and some conventions were inherited from Torch 7's NN and cuNN libraries.
(EDIT: and the name "autograd" originates from HIPS autograd library, which I think predates Chainer)
NubaCUDA gave me lots of small problems and a few big ones. The poor support for debug/perf tools and poor integration with other high-level python CUDA code (FFTs in particular) sent me packing, but the number of small problems was excessive in comparison to the size of my code. I had 5 reduced bugs at the bottom of my notebook and two paragraphs of "baggage" at the top to support a tiny little 50LoC kernel: one paragraph for the environment variables and one for patching nubacuda itself for a trivial API incompatibility that hadn't been fixed for the better part of a year. All of this for a tool that provided a diminutive subset of functionality at the intersection of both python and C. I've felt more computational freedom writing BASIC on my TI-83.
CuPy could well have changed that equation!
> incomprehensible error messages when the type inference goes wrong
NumbaCUDA is truly the galaxy-brain of type checking: first it complains loudly so as to force you to provide type information, then it opts to not complain about a mismatch, and then it silently reinterpret_casts a double* to float* behind your back.
I know it's free software and I have no right to complain, but I sure sunk a lot of time into this dead end and regret it.
Spiffy icon though.
If you’re doing custom kernels you should take a look at the Julia library CuArray [1] and generic kernels [2]. I really like that I don’t have to dig into C++ and deal with all of the memory and kernel management.
1: https://github.com/JuliaGPU/CuArrays.jl 2: http://mikeinnes.github.io/2017/08/24/cudanative.html
I love Julia, but I haven't managed to convert anyone else on my team and I already spent my informal exploration budget for the GPU project on nubacuda, so JuliaGPU will have to wait for another time. I'll be sure to keep it in mind, though!
How is the CUDA debug/perf story with Julia? Does it play nice with the nvidia tooling?
I haven't dug too deep with CudaNative / Cuarray to understand the state of Julia perf debugging. Though here's one post on the topic:
https://discourse.julialang.org/t/cudanative-is-awesome/1786...
In general It's been very pleasant experimenting with gpu programming in Julia. I couldn't quite grok tensorflow code, and it's cool to just declare a Julia array and send it the GPU.
It takes a couple of microseconds just to start a kernel much less the time it takes to transfer data back and forth.
I've used PyCUDA for a very long time. You can use Cython/ctypes/cffi to pass PyCUDA arrays to standard C/CUDA code.
https://pytorch.org/get-started/locally/
E.g "To install PyTorch via Anaconda, and you are using CUDA 9.0, use the following conda command". If they are shipping with CUDA perhaps that should be phrased more like "and you want to use CUDA 9.0". And of course you do indeed need your own CUDA installation if you want to build PyTorch from source yourself.
[0] https://tomaugspurger.github.io/pandas-extension-arrays.html