That is a bold claim. It seems to me they simply assume all necessary code would trivially be converted to run on GPUs.
That is a bold claim. It seems to me they simply assume all necessary code would trivially be converted to run on GPUs.
Also, we tried to get into NVIDIA Inception and were rejected for the company being too small, so even if we wanted to do GPU deployment, we wouldn't be able to buy 3090s for testing. Scalpers / eBay are not an option in Germany because for company purchases, I need a proper tax invoice from another VAT-registered company.
You can start with things like switching the weights and activations to fp16. A100 also supports BFloat16, which is even better and should work right out of the box.
At the time, the price of the A4000 was similar to scalper-prices for similarly-performing gaming cards, but:
- Had ECC RAM and was generally designed for ML/enterprise/data workloads
- Was more power-efficient
- Had 16GB RAM, which you won't get on anything short of an A3090
I'd taking a consumer card at MSRP over the A4000, but not at street price. For what I'm doing, an A3060 Ti would probably even have been fine. It has an MSRP of $330, but a street price of $500-$1000 depending on the month you look. The enterprise cards tend to sell at MSRP (which is just around a grand for the A4000).
Note that NVidia reuses part numbers. With "4000," there is also an older NVIDIA Quadro RTX 4000 (no A) selling for the same price on eBay / Amazon / etc. as the Ampere-based A4000.
The biggest hurdles are learning the dev tech stack and setting up the development environment.
The last time I tried out CUDA, it took me a couple of days to correctly setup the environment; just to spend an hour playing around parallelizing math operations.
The best way I could describe the first impression dev experience: It’s a huge hack, adding non-standard language extensions to C/C++
And so relying on their own compiler or linker (iirc), instead of the standard C++.
Couple this with NVidia not playing nice with OpenCL to force you into their playground, and I lost interest pretty quickly.
There are cases where the random access to RAM are very deep in the algorithm. Consider e.g. sparse matrix multiplication.
Some easy-to-parallelize math and vector operations can yield impressive results without it being anywhere near optimal.
I do agree that optimal requires mastery of GPU programming; but I also think GPU programming can yield good results without said mastery, and be quite accessible
If anything it should be easier than writing raw multi-threading code, because GPUs aren’t (weren’t?) general purpose calculation units.
So more bare metal concerns, but less to reason about overall.
GPU instances being absurdly pricey is not a good enough reason?
The hangups mostly seem to be in that a lot of scientific code hasn't been written with GPU style parallelization in mind, so updating the code takes time. That said, since a lot of newer supercomputers are including a lot of GPUs, there is a lot more incentive for developers to attempt to port at least the parts that can benefit from GPUs.
From my (admittedly limited) experience, the biggest hangup is convincing scientists that GPUs aren't just another tech fad and it's worth the effort to port. There also seems to be a bit of uncertainty over which platform to trust, OpenCL support has been shaky on most vendors and is fairly primitive in terms of features, CUDA is tied to one company and AMD doesn't have a great record of supporting HIP across their product stack, so choosing any one is a big decision (comparable to choosing to use C/C++ or Python, which are effectively guaranteed to be supported on every supercomputer for the next decade or two).
Of course, matrix algebra was only one of the problems. BLAS supports the writing of LAPACK routines, which tend to deal more with dense factorizations and eigenvalue problems. I believe more factorizations have been implemented recently, but I'm not currently up to date. Nevertheless, that was absolutely a bottle neck for the longest time. Yes, iterative methods don't need factorizations, but a good fraction of the preconditioners do.
Then, of course, this speaks nothing of the sparse linear algebra problem. There's multiple ways to do it, but many of the good sparse factorization routines need dense factorization routines, so these needed to come onto the market first and that took time.
And, to be clear, I know that there are multiple, good teams working on this. It takes time. And, there's a huge number of operations. If you're bored, go look at the manual for Intel's MKL and see the number and variety of operations that it provides. Those operations are there because people like me need them to do our job. I'll also agree that the kinds of operations we use to do our job will evolve over time with hardware. However, matrix algebra, factorizations, and eigenvalue problems lie at the very core of applied mathematics and expecting the mathematics to rework the last several hundred years of practice to accommodate the lack of tooling from the GPU manufacturers isn't realistic either.
Anyway, if someone knows the current state of what's possible, I'd love to hear. Selfishly, what it really boils down to is what operations (algebra, factorization, or eigenvalue), how big (how much memory or on multiple GPUs), and dense or sparse.
Personally, I don't really care where the libraries come from, but I will contend that the hardware manufacturers are generally the best place for this work to occur. For many years, each of the chip makers published their own high performance BLAS and LAPACK routines, which worked really well. Intel had MKL, AMD had AMCL, Sun had sunperf, IBM had ESSL. NVIDIA does the same, and I'm hugely grateful for that, but it's not complete.
Really, though, the comment is more to answer why more mathematicians don't use GPUs. My contention is not that I or my colleagues view it as a fad, but more that there's a lack of routines that we depend on.
I currently work for myself, but I'll also mention that the manufacturers did go to management at places I've worked and management did put pressure on staff to just rewrite everything in CUDA. As staff does, they said, ok, fine, but it's going to cost you labor. Management didn't have the money, or didn't want to spend it, and some mild office conflict occurred.
Anyway, mostly that's to say that I will gladly spend thousands on GPUs and recommend it to my clients as soon as I don't have to write all of the low level routines myself.
How much of a speed up would make it worth your time to implement it yourself?
CUDA is 15 years old. All top supercomputers use GPUs and have been for a while. Mature software tools like cupy exist to make GPU programming easier. I don’t think anyone thinks today “GPUs are a fad”. The problem is a lot of code is simply hard to parallelize, and would need to be rewritten from scratch, likely in another language. For large old codebases this is a massive effort.
This was what led to the hesitance from the people I worked with. They were unsure if in another generation or two they'd have to look at another large effort to support another new programming paradigm. There was also hesitance in relying on CUDA since it locks them into a vendor.
From an engineer's point of view, yeah, GPUs obviously aren't a fad, but a lot of this software is written by scientists, where the computer is simply a means of getting their result and not something they keep up with as they would with their own field and yet the software needs to remain relatively stable to be reliable for research. Thus anyone asking for a large modification of the code is going to be viewed with extreme skepticism.
For reference, I had to spend around 2 months of weekly meetings and presentations showing test results and discussing the risk/reward tradeoffs in detail to convince a group that a limited port of the code would be worth going ahead with and that CUDA was effectively the only reliable option for now. They've only gotten serious about a more in-depth port from seeing the large speedups without breaking compatibility which we managed to get after a few iterations.
Deep ML is about the only thing that's both large and "classic parallelizable HPC".