For like 8 years their drivers on Linux were a nightmare and AMD could have come in and done better.
For like 8 years their drivers on Linux were a nightmare and AMD could have come in and done better.
AMD eventually did while Nvidia's drivers remained a nightmare almost until these days. But sure, AMD could have done it sooner.
and yet that trillion-dollar valuation built over the last decade is built with customers almost entirely running on those "nightmare" linux drivers, while AMD's linux drivers crash running the sample app on supported hardware+OS, and nobody at AMD cared until finally a tech-bro with a loud enough platform shamed them into fixing it...
... and this is something like AMD's third crack at the apple, and the first three sets of drivers (one of which is literally a Vulkan-branded spec) are just as non-functional today as rocm was a year ago.
(OpenCL, Fusion HSA/AMD APP, Vulkan Compute/SPIR-V... all still broken so badly that Octane called them out for being unable to successfully compile their renderer and for lack of vendor support, so badly that Blender pulled support after years of turbulent and poorly-performing attempts to work with AMD, etc)
Hell, nvidia drivers might be often complained about, but for years I would take nvidia because the crappiness was manageable and close to nothing if you were in the target market (desktop workstations running X11 on only nvidia GPUs? The only issue was if you were running super latest kernel).
Now that tools like Blender and the like are increasingly picking up Vulkan support, there is no reason for the above to use Nvidia anymore.
It was when you went outside said use case that things started getting worse, and you had to wait long time for fixes. Sometimes it was because the changes in XFree/X.Org were effectively fixated on how some other vendors did things (cough intel cough), or involved things that effectively nobody wanted to spend engineering to fix properly (like rebuilding rendering path to be able to handle hybrid graphics properly when hybrid graphics came into world years after critical set of X.Org devs decided to stop any real development into X.Org...).
Vulkan Compute also is nowhere close to feature parity with CUDA, so not sure it would be picked up instead.
Also, it's not like Wayland offered a concrete target to support, different approaches of how to actually provide device context to applications were from beginning a "now draw the rest of the owl" thing.
Nowadays almost nobody cares about OpenCL.
"...no, but you could expand on OpenCL or Vulkan compute if you wanted. There are other spec stakeholders, we can't give you carte-blanche control, Apple."
"Why do you insist upon mismanaging the industry's APIs? Screw you guys!" <Beginning of mid 2010s "Khronos Drought" at Apple Computers>
Like, I started using CUDA (through frameworks) over ten years ago, and basically nobody has come up with anything competitive since then.
This is a significant understatement. For quite some time Jensen has been saying repeatedly that 30% of their R&D spend is on software. With the money-printing machine that is Nvidia if that holds they're going to continue to rocket ahead of competitors in terms of delivering actual solutions.
The "What are you talking about? AMD/Intel runs torch just fine!" crowd clearly haven't seen things like RIVA, Deepstream, Nemo, Triton Inference Server/NIM, etc. Meanwhile AMD (ROCm) still struggles with flash attention...
What these hardware-first (only?) companies like AMD don't seem to understand is that people buy solutions, not GPUs. It just so happens that GPUs are the best way to run these kinds of workloads but if you don't have a wholistic and exhaustive overall ecosystem you end up in single digit market share vs Nvidia at ~90%.
"What are you talking about? AMD/Intel runs torch just fine!" refers indirectly to the value of having competition in markets, not jump on the (well-funded,slick) monopoly bandwagon.
The experience was incredibly simple: write C like usual but annotate a few C functions with some extra keywords and compile using a custom frontend/preprocessor/whatever-nvcc-was instead of gcc (i was on Linux - and BTW i heavily contest the notion that Nvidia drivers on Linux were "nightmare", they always worked just fine with both performance and features comparable to their Windows counterparts while ATi/AMD had buggy and broken drivers for years). Again, the experience was very simple, i even just copy/pasted a bunch of existing C code i had and it worked.
Later i tried to use OpenCL which was supposedly the open alternative. That one felt way more primitive and low level, like writing shaders without the shading bits.
In a way, as you wrote, it was kinda like DirectX: that is, CUDA was like using OpenGL 1.1 with its convenient and straightforward C API and OpenCL was like using DirectX 3 with its COM infested execute buffer nonsense.
After that i never really used CUDA (or OpenCL for that matter) but it gave me the impression that Nvidia did put way more effort on developer experience.
Since CUDA 3.0, NVidia has embraced a polyglot stack, with C, C++ and Fortran at the center, and PTX for anyone else.
Followed by changing CUDA memory model to map that of C++11.
Khronos never cared for Fortran, and only designed SPIR, when it became obvious they were too late to the party.
So not only has CUDA first level tooling for C, C++, Fortran, with IDE integration in Visual Studio and Eclipse, graphical GPU debugger with all the goodies of a modern debugger, it also welcomes any compiler toolchain that wants to target PTX.
Java, Haskell, .NET, Julia, Python JITs, .... there are plenty to chose from, without going through "compile to OpenCL C99" alternative.
Finally, the myriad of libraries to chose from.
CUDA is not only for AI, by the way.
And because of that, their OpenCL implementation also works better than others. So there's more tooling not just from nvidia using it, because it. just. works.
Compare this with AMD, whose latest framework is a total mess of "will it work on this GPU?", sometimes needing custom wrangling to enable, etc. etc. and it's effectively supported only on the most expensive compute-only cards.