They were a key cause for LLMs being a thing in the first place.
NVidia has continued to stay ahead because every alternative to CUDA is half baked trash, even when the silicone make sense. As a tech company, trading $$ for time not spent dealing with compatibility bugs and broken drivers pretty much always makes business sense.
> dealing with compatibility bugs
> broken drivers
Describes my experience trying to use CUDA perfectly.
We have a long way to go and we haven't even started yet.
If you need people to abandon an ecosystem thats been developed steadily over nearly 20 years for your shiny new thing in order to keep it around, you'll never compete.
But AMD and others could have done the same, had they been better at riding that wave. There’s a reason it was Nvidia who won out.
I was using it as soon as it came out in 2007 and my dinky desktop workstation was out performing the main frame in the basement.
Twenty years ago I was thinking we'd be speccing machines in kilocores by now.
> Number of SMs is a more appropriate equivalent to CPU core count.
What do you mean by this? Why should an SM be considered equivalent to a CPU core? An SM can do 128 simultaneous adds and/or multiplies in a single cycle, where a CPU core can do, what, 2 or maybe 4? Obviously depends on the CPU / core / hyperthreading / # math pipelines / etc., but the SM to CPU-core ratio of the number of simultaneous calculation is in the double digits. It’s a tradeoff where the GPU has some restrictions in return for being able to do many multiples more at the same time.
If you consider an SM and a CPU equivalent, then the SM’s perf can exceed the CPU core by ~2 orders of magnitude — is that the comparison you want? If you consider a GPU thread lane and a CPU thread lane equivalent, the the GPU thread lane is slower and more restricted. Neither comparison is apples to apples, CPUs and GPUs are made for different workloads, but arguing that an SM is equivalent to a CPU core seems equally or more “misleading” when you’re leaving out the tradeoff.
I’d argue that comparing SMs to cores is misleading and that it makes more sense to compare chips is by their thread counts. Or, don’t compare cores at all and just look at the performance in, say, FLOPS.
That's using SIMD, but so is Nvidia for all intents and purposes. Those "cuda cores" aren't truly independent: when their execution diverges, masking is used pretty much like you'd do in CPU SIMD.
A lot of the control logic is per-SM or perhaps per-SIMD unit -- there are multiple of those per SM. You could perhaps make a case that it's the individual SIMDs which correspond to CPU cores (that makes the flops line up even more closely). It depends on what the goal of the comparison is.
https://images.nvidia.com/aem-dam/Solutions/geforce/news/rtx...
An SM is split into four identical blocks, and I would say each block is roughly equivalent to a CPU core. It has a scheduler, registers, 32 ALUs or FPUs, and some other stuff.
A CPU core with two AVX-512 units can do several integer operations plus 32 single-precision operations (including FMA) per cycle. Not 2 or 4. An older CPU with 2-3 AVX2 units could fall slightly behind, but it's pretty close.
That doesn't factor in the tensor units, but they're less general purpose, and CPUs usually put such things outside the cores.
I would say an SM is roughly equivalent to four CPU cores.
Yes, considering CPU SIMD, maybe comparing a CPU core to a CUDA warp makes some sense in some situations. The peak FLOPS rate is still so much higher on Nvidia though, that the comparison hardly makes sense. So yeah like I and the other commenter mentioned, it depends entirely on what comparison is being made.
Nvidia/CUDA process: Download package. Run the build. It works. Run your thing- it's GPU accelerated. Go get a beer/coffee/whatever while your net runs.
AMD process: Download package. Run the build. Debug failure. Read lots of articles about which patches to apply. Apply the patches. Run the build. It fails again. Shit. OK ok now I know what to do. I need a special fork of the package. go get that. Find it doesn't actually have the same API that the latest pytorch/tf relies on. OK downgrade those to an earlier version. OK now we're good. Run the build again. Aw shit. That failed again. More web searches. Oh ok now I know - there's some other patches you need to apply to this branch. OK cool. Now it compiles. Install the package. Run the thing. Huh. That's weird. GPU accelleration isn't on.... sigh....
- CUDA conception in 2006 to build super computers for scientific computing.
- CUDA influencing CPU designs to be dual purpose, with major distinction of RAM amounts (for scientific compute you need a lot more RAM compared to gaming)
- Crypto craze driving extreme consumer GPU demand which enabled them to invest heavily into RND and scale up production.
- AI workload explosion arriving right as the crypto demand was dying down.
- Consistently great execution, or at least not making any major blunders, during all of the above.
It doesn’t mean they didn’t make a bunch of mistakes its that when they did there was no competition to to realistically turn towards, and they fixed a lot of their mistakes.
Management at the two could not be more opposite if they tried.
Nvidia's advantage is that they have by far the most complete programming ecosystem for them. (Also honestly… they're a meme stock.)
The first transformer models were developed at Google. NVIDIA were the card du jour for accelerating it in the years since and have contributed research too, but your statement goes way too far
The first transformer models didn’t even use CUDA and CUDA didn’t have mass ecosystem inroads till years later.
I’m not trying to downplay NVIDIA but they specifically mentioned cause and effect, and then said it was because of NVIDIA.
NVIDIA absolutely contributed to the foundation but they are not the foundation alone.
Alexnet was great research but they could have done the same on other vendors at the time too. The hardware didn’t exist in a vacuum.
To say it wouldn’t have been possible on AMD is ludicrous and there is a pattern to your comments where you dismiss any other companies efforts or capabilities, but are quite happy to lay all the laurels on NVIDIA.
The reality is that multiple companies and individuals got us to where we are, and multiple products could have done the same. That's not to take away from NVIDIA's success, it's well earned, but if you took them out of the equation, there's nothing that would have prevented the tech existing.
> It was just what happened to be available and most straightforward at the time
AMD made better hardware for a while and people wanted OpenCL to succeed. The reason why nvidia became dominant was because their competitors simply weren’t good enough for general purpose parallel compute.
Would AI still have happened without CUDA? Almost certainly. However nvidia still had a massive role in shaping what it looks like today.
It’s one of the projects I think the Khronos group mishandled the most unfortunately.
If I were to categorize the successful ones:
glTF, KTX, SPIR-V, OpenGL and its variants, WebGL
People will say Vulkan but it has the same level of adoption as OpenCL, and has the same issue that it competes against vendor specific APIs (DX and Metal) that are just better to use. It’s still used though of course as a translation target but imho that doesn’t qualify it as a success.
OpenCL was and is a failure of grand magnitude. As was colada.
I graduated college in 2010 and I took a class taught in CUDA before graduating. CUDA was a primary driver of NN research at the time. Sure, other tools were available, but CUDA allowed people to build and distribute actually useful software which further encouraged the space.
Could things have happened without it? Yeah, for sure, but it would have taken a good deal longer.
Deep learning was only possible because you could do it on NVidia cards with Cuda without having to use the big machines in the basement.
Trying to convince anyone that neural networks could be useful in 2009 was impossible - I got a grant declined and my PhD supervisor told me to drop the useless tech and focus on something better like support vector machines.
The difference is AMD killed theirs to use OpenCL and NVIDIA kept CUDA around as well.
I tried using AMD Stream and it lacked documentation, debugging information and most of the tools needed to get anything done without a large team of experts. NVidia by comparison could - and was - used by single grad students on their franken-stations which we build out of gaming GPUs.
The less we talk about the disaster that the move to opencl was the better.
That said ROCm is quite a recent thing borne out of acknowledging that OpenCL 2 was a disaster.
OpenCL 1 had a reasonable shot and was gaining adoption but 2 scuppered it.
I did write quite a bit of OpenCL prior to that on Intel/AMd/NVIDIA, both for training and for general rendering though, and did some work with Stream before then.
Cuda by comparison JustWorks^tm.
1 was definitely a lot easier to work with than 2. CUDA is easier than both but I don’t think I hit anything I could do in CUDA that I couldn’t do in OpenCL, though CUDA of course had a larger ecosystem of existing libraries.
Thanks for the trip down memory lane.
It could have been AMD/ATI profiting from such random events, as they were the ones that financed AI development.