They were a key cause for LLMs being a thing in the first place.
They were a key cause for LLMs being a thing in the first place.
NVidia has continued to stay ahead because every alternative to CUDA is half baked trash, even when the silicone make sense. As a tech company, trading $$ for time not spent dealing with compatibility bugs and broken drivers pretty much always makes business sense.
> dealing with compatibility bugs
> broken drivers
Describes my experience trying to use CUDA perfectly.
We have a long way to go and we haven't even started yet.
If you need people to abandon an ecosystem thats been developed steadily over nearly 20 years for your shiny new thing in order to keep it around, you'll never compete.
But AMD and others could have done the same, had they been better at riding that wave. There’s a reason it was Nvidia who won out.
I was using it as soon as it came out in 2007 and my dinky desktop workstation was out performing the main frame in the basement.
Twenty years ago I was thinking we'd be speccing machines in kilocores by now.
> Number of SMs is a more appropriate equivalent to CPU core count.
What do you mean by this? Why should an SM be considered equivalent to a CPU core? An SM can do 128 simultaneous adds and/or multiplies in a single cycle, where a CPU core can do, what, 2 or maybe 4? Obviously depends on the CPU / core / hyperthreading / # math pipelines / etc., but the SM to CPU-core ratio of the number of simultaneous calculation is in the double digits. It’s a tradeoff where the GPU has some restrictions in return for being able to do many multiples more at the same time.
If you consider an SM and a CPU equivalent, then the SM’s perf can exceed the CPU core by ~2 orders of magnitude — is that the comparison you want? If you consider a GPU thread lane and a CPU thread lane equivalent, the the GPU thread lane is slower and more restricted. Neither comparison is apples to apples, CPUs and GPUs are made for different workloads, but arguing that an SM is equivalent to a CPU core seems equally or more “misleading” when you’re leaving out the tradeoff.
I’d argue that comparing SMs to cores is misleading and that it makes more sense to compare chips is by their thread counts. Or, don’t compare cores at all and just look at the performance in, say, FLOPS.
That's using SIMD, but so is Nvidia for all intents and purposes. Those "cuda cores" aren't truly independent: when their execution diverges, masking is used pretty much like you'd do in CPU SIMD.
A lot of the control logic is per-SM or perhaps per-SIMD unit -- there are multiple of those per SM. You could perhaps make a case that it's the individual SIMDs which correspond to CPU cores (that makes the flops line up even more closely). It depends on what the goal of the comparison is.
https://images.nvidia.com/aem-dam/Solutions/geforce/news/rtx...
An SM is split into four identical blocks, and I would say each block is roughly equivalent to a CPU core. It has a scheduler, registers, 32 ALUs or FPUs, and some other stuff.
A CPU core with two AVX-512 units can do several integer operations plus 32 single-precision operations (including FMA) per cycle. Not 2 or 4. An older CPU with 2-3 AVX2 units could fall slightly behind, but it's pretty close.
That doesn't factor in the tensor units, but they're less general purpose, and CPUs usually put such things outside the cores.
I would say an SM is roughly equivalent to four CPU cores.
Yes, considering CPU SIMD, maybe comparing a CPU core to a CUDA warp makes some sense in some situations. The peak FLOPS rate is still so much higher on Nvidia though, that the comparison hardly makes sense. So yeah like I and the other commenter mentioned, it depends entirely on what comparison is being made.
Nvidia/CUDA process: Download package. Run the build. It works. Run your thing- it's GPU accelerated. Go get a beer/coffee/whatever while your net runs.
AMD process: Download package. Run the build. Debug failure. Read lots of articles about which patches to apply. Apply the patches. Run the build. It fails again. Shit. OK ok now I know what to do. I need a special fork of the package. go get that. Find it doesn't actually have the same API that the latest pytorch/tf relies on. OK downgrade those to an earlier version. OK now we're good. Run the build again. Aw shit. That failed again. More web searches. Oh ok now I know - there's some other patches you need to apply to this branch. OK cool. Now it compiles. Install the package. Run the thing. Huh. That's weird. GPU accelleration isn't on.... sigh....