It's hard to find skills that don't have a degree of provincialism. It's not a great feeling, but you more on. IMO, don't over-idealize the concept of general-knowledge to your detriment.
I think we can also untangle the open-source part from the general/provincial. There is more to the world worth exploring.
Single socket 8 core CPU? Yes.
If you spent some time playing with trying to eke out performance on Xeon Phi and have done NUMA-aware code for multi socket boards and optimising for the memory hierarchy of L1/L2/L3 then it really isn't that different.
You move from one thing to the next.
With your transferable skills, experience and thinking that is beyond one programming language.
Even Apple is simply exporting to CUDA now.
Really!!! Any resources you can share?
https://9to5mac.com/2025/07/15/apples-machine-learning-frame...
This is like when journalists write clickbait article titles by omitting all qualifiers (eg "states banning fluoride" when it's only some states).
One framework added a CUDA backend. You think all of Apple uses only one framework? Further what makes you think this even gets internal use?
Only that Apple not only might use CUDA internally but made a public release available too.
CUDA seems to be a trigger word in this thread for some.
https://9to5mac.com/2025/07/15/apples-machine-learning-frame...
what does this sentence mean?
> Apple is simply exporting to CUDA now.
There are plenty of acceptable styles. The guidelines don't insist on only one style.
Stories of exploring DOS often ended up at hex editing and assembly.
Best to learn with whatever options are accessible, plenty is transferable.
While there is an embarrassment of options to learn from or with today the greatest gaffe can be overlooking learning.
ROCm is getting some adoption, especially as some of the world's largest public supercomputers have AMD GPUs.
Some of this is also being solved by working at a different abstraction layer; you can sometimes be ignorant to the hardware you're running on with PyTorch. It's still leaky, but it's something.
I used to use ROCFFT as an example, it was missing core functionality that cuFFT has had since like 2008. It looks like they've finally caught up now, but that's one library among many.
Programming languages are groups of syntax.
There are enough people for who it's worth it, even if just for tinkering, and I'm sure you are aware of that.
It reads a bit like "You shouldn't use it because..."
Learning about Nvidia GPUs will teach you a lot about other GPUs as well, and there are a lot of tutorials about the former, so why not use it if it interests you?
Just some fundamentals I can think of off the top of my head. I'm surprised people saying that the lower level systems/hardware stuff are untransferable. These things are used everywhere. If anything, it's the AI itself that's potentially a bubble, but the fundamental need for understanding performance of systems & design is always there.
Re Oracle and "big 90s names" specifically, there is a lot of it out there. Maybe it never shows up in the code interfaces HNers have to exercise in their day jobs, but the tech, for better or worse, is massively prevalent in the everyday world of transit systems and payroll and payment...ie all the unsexy parts of modern life.
And wait until I tell you about my Cobol open seats - on modern Linux on cloud VMs too! :-)
The software is proprietary, and easy to ignore if you don't plan to write low-level optimizations for NVIDIA.
However, the hardware architecture is worth knowing. All GPUs work roughly the same way (especially on the compute side), and the CUDA architecture is still fundamentally the same as it was in 2007 (just with more of everything).
It dictates how shader languages and GPU abstractions work, regardless of whether you're using proprietary or open implementations. It's very helpful to understand peculiarities of thread scheduling, warps, different levels of private/shared memory, etc. There's a ridiculous amount of computing power available if you can make your algorithms fit the execution model.
Sounds good on paper but unfortunately I've had numerous issues with these "abstractors". For example, PyTorch had serious problems on Apple Silicon even though technically it should "just work" by hiding the implementation details.
In reality, what ends up happening is that some features in JAX, PyTorch, etc. are designed with CUDA in mind, and Apple Silicon is an afterthought.
Work keeps us humble enough to be open to learn.
When cuda rose to prominence were there any viable alternatives?
Better not learn CUDA then.
For most IT folks it doesn't make much sense.