DeepSeek's AI breakthrough bypasses industry-standard CUDA, uses PTX
tomshardware.com
tomshardware.com
Well, it was quite a reverse culture shock after I moved to the US. I definitely didn't know that "teacher's pet" was a thing, or my coworker, a brilliant engineer who went to a highly reputed public school, was chased off his school bus simply because he used some poetic words, or geeks were not that respected in schools, or a mile wide and an inch deep with great leadership is what the US people revered. In the meantime, I guess other countries more or less picked up the baton of the US culture, and grew their own geeks.
This is the first time it’s been so bad since, but I’m optimistic: we’re actually in a better position now because serious foreign tech dynamos don’t give a fuck about American mafia nepotism, we can’t just keep this in the family.
I don’t know if this DeepSeek App Store thing will be the match that lights this thing up, but the grass is very tall and very, very dry.
as an adult across the pond, I'm very disappointed by the US.
There exist quite a lot of US-American movies. I would claim that those who are less deeply centered around US-specific cultural traits are typically much more revered in other countries.
There are, of course, exceptions from this rule (for example "The Simpsons", which excessively satirizes the USA's popular culture, but is nevertheless loved in many other countries), but I do think that the general rule of thumb does hold.
In this sense: even if you watch US-American movies, it is rather easy to mostly ignore those who strongly show US-specific cultural traits and habits, in particular if these are considered to be annoying in the other country.
The lack of a PhD probably has more to do with the structure of Chinese education but these people were basically studying like PHDs.
https://www.firstpost.com/explainers/china-deepseek-ai-full-...
China has a reputation for being a "paper mill" place though, so not sure that publishing papers has any value whatsoever as a quality indicator.
Only those tech talents in the USA who are useful for some big tech agenda are not languishing.
https://www.bloomberg.com/graphics/2023-black-lives-matter-e...
Job growth amongst foreign born residents in the US has significantly outpaced native born job growth:
https://www.bls.gov/opub/ted/2023/foreign-born-workers-were-...
Large portions of the world are anti-intellectual, at the same time intellectuals are often much worse than the average person and frequently do deserve scorn as a class of people in society.
CUDA is not an industry standard. Vulcan is an industry standard. They did not bypass CUDA... that's like saying if I use Vulcan I'm bypassing OpenGL. PTX is an alternative low level API provided by Nvidia because of how awful CUDA is for high performance code.
What DeepSeek wrote could only have either been written in PTX or Vulcan.
Any other company could have done this, and low latency traders on Wall Street that use Nvidia write their stuff in PTX for obvious reasons.
OpenAI, was, is, and always will be, absolutely incompetent when it comes to using their hardware effectively... and they're no different than any other company. Reading is not a goddamned super power! Just read the docs!
We do a lot of forecasting and solvers where I am, just run them on CPUs though.. but maybe if you’re wanting to compete on speed you would?
This depends a lot on the problem and the algorithm that is used. For example interior point methods are clearly better suited to be running on GPUs than the primal or dual simplex algorithm.
PTX is not the IL used by Nvidia's drivers, but does compile directly to it with less slop involved. If you had said "PTX's instructions are analogous to writing assembly for CPUs or any other GPUs (ala Clang's AMDGPU target)", that would have probably been the better way.
Arguably, PTX is closer to being the SPIR-V part of their stack (more than just an assembler compiler, but similar in concept). None of Nvidia's tools really ever line up with good analogies with the outside world, the curse of Nvidia's NIH syndrome.
Generally, you're not going to be writing all of your code in PTX, but I find it wild you think people going to "an average technical university" would be unable to use it for the parts they need it for. That says more about you than it does them.
All of Nvidia's docs for this are online, it isn't that hard. Have you tried?
How else would you have understood it? At this level it's literally just pedantics. In the same way you can say C doesn't technically compile to assembly for CPUs. The point is that it's the lower abstraction level that is still (more or less) human readable. But just like in CUDA, you may want to write parts of your code in it if you want to benefit from things that the higher level language doesn't expose. The terminology might seem different, but in practice it is pretty analogous.
With PTX, you are not writing in the native assembly of the Nvidia GPU, but yet another abstraction, just one more similar in nature.
I wonder if you vould you point me to concrete examples where people write PTX rather than CUDA? I'm asking because I just learned CUDA since it's so much faster than Python!
Open source authors typically shy away from Nvidia closed source APIs, and PTX is tied to how Nvidia hardware works, so you won't see it implemented for other hardware.
To do what Deepseek did, but didn't want to waste your time and money with Nvidia, you'd use Vulkan. Theres more Vulkan in the world than CUDA.
For various micro-bench reasons I wanted to use a global clock instead of an SM-local one, and I believe this was needed. Also note that even CUDA has "lower level"-like operations, e.g. warp primitives. PTX itself is super easy to embed in it like asm.
> as long as they have no real competitors in that space.
?! isn't _avoiding competition_ a good enough incentive for them to improve their offering?
It’s been years, and they’re doing… what exactly to catch up?
I'd also bet that if you buy enough of the Nvidia chips, they'll probably send in a bunch of engineers to get everything working with their full stack. AMD won't be able to do that in the same way because they're not vertically integrated.
Yep. The interconnect is the most overlooked part of the stack.
aren’t APIs no longer copyrightable? they should just have reimplemented cuda with their backend
> Specifically, we employ customized PTX (Parallel Thread Execution) instructions and auto-tune the communication chunk size, which significantly reduces the use of the L2 cache and the interference to other SMs.
So they have some intrinsic in some part of their training framework. That's it.
there's no finally - perf people have always written kernels in asm.
IIRC this is still relatively hardware agnostic. Can you actually get very far by doing this ? From a quick perusal, DeepSeek also uses Triton in the codebase.
Isn't CUDA an Nvidia child ? This sounds like "Microsoft = industry standard".
This gives them a few months head start before meta and Google start doing the same thing.