Jim Keller criticizes Nvidia's CUDA, x86
tomshardware.com
tomshardware.com
> Even Nvidia itself has tools that do not exclusively rely on CUDA. For example, Triton Inference Server is an open-source tool by Nvidia that simplifies deploying AI models at scale, supporting frameworks like TensorFlow, PyTorch, and ONNX. Triton also provides features like model versioning, multi-model serving, and concurrent model execution to optimize the utilization of GPU and CPU resources.
> Nvidia's TensorRT is a high-performance deep learning inference optimizer and runtime library that accelerates deep learning inference on Nvidia GPUs. [...]
Keller was speaking of OpenAI's Triton (https://openai.com/research/triton), a Python-like language that is compiled to code for Nvidia GPUs, but Tom's Hardware mixed this up with Nvidia's Triton Inference Server, a higher level tool that's really not a replacement for CUDA and not directly related to the Triton language. Easy to confuse these if you are a writer in a hurry.
Wasn't this one of the reasons AMD abandoned/deprioritized their efforts on such a project?
> His statements also imply that even though he has worked stints at some of the largest chipmakers in the world, including the likes of Apple, Intel, AMD, Broadcom (and now Tenstorrent), we might not see his name on the Nvidia roster any time soon.
According to the video, Keller’s philosophy emphasizes the importance of both theory and engineering in the field of computer science. He believes that theory provides a foundation for understanding how things work, while engineering is the practical application of that knowledge. He argues that both are essential for making progress in the field.
Keller also emphasizes the importance of craftsmanship and attention to detail. He believes that the best engineers are those who take pride in their work and are constantly striving to improve it. He believes that this is essential for building high-quality, reliable computer systems.
Finally, Keller believes that it is important to be open to new ideas and to be willing to experiment. He believes that this is the best way to make progress in the field of computer science.
Here are some specific examples from the podcast that support Keller's philosophy:
- Keller discusses the importance of theory in the development of branch prediction, a key technique for improving computer performance. He explains that while the basic idea behind branch prediction was known for many years, it was only through theoretical advances that it was possible to develop a practical implementation.
- Keller also discusses the importance of engineering in the development of the Alpha 21264 microprocessor. He explains that while the chip was a groundbreaking design, it also had some flaws that were only discovered after it was released. He says that these flaws could have been avoided if the engineers had paid more attention to detail.
- Finally, Keller discusses the importance of being open to new ideas. He talks about his work on the TenstorFlow chip, which is a new type of chip designed for machine learning applications. He says that he was initially skeptical of the idea, but that he eventually came to believe that it had the potential to be a major breakthrough.
Overall, Keller's philosophy is one of pragmatism and open-mindedness. He believes that the best way to make progress in computer science is to be willing to experiment and to learn from both successes and failures.
...
Here are some timestamps related to Keller's philosophy in the YouTube video:
1:18:02 - Keller discusses the importance of both theory and engineering in good design.
1:23:22 - Keller gives the example of branch prediction as a breakthrough in engineering that was based on theory.
1:34:12 - Keller talks about the importance of craftsmanship in engineering.
1:42:15 - Keller discusses the limitations of human thinking and the importance of being open to new ideas.
2:12:22 - Keller talks about the responsibility of engineers to society.
Or, as is just often the case, he competes with Nvidia because he has a different opinion (of their vision).
How does someone achieve this? Raw IQ? The right schools? Right place/right time? Luck + effort?
Is there something a moderately intelligent person can do that has a high likelihood of moving them in this direction (if not actually achieving it)?
The field itself is already pretty niche. The amount of people on his level probably fit on an index card. That helps probably the most.
He's obviously a smart guy, but to me his wisdom comes from being able to actually understand the system and explain it. Whether he actually codes it or is even that smart becomes irrelevant, he has the ability to big picture it and get aces in their places.
But mostly the first point. The niche is so small you'd probably be hard pressed to find someone in comments who actually has worked alongside him to even attest to any of this.
It's immensely clear that when he goes somewhere, some big advancement usually happens, and he can explain why. Even if that's -all- he knew, it'd be enough to make him super valuable. So if we have to keep going back and pinning traits, I'd say it's some amount of intelligence mixed with a whole lot of passion.
More and more I'm doing more branchy algorithms so the CPU is now my new bottleneck - thank you fast GPUs - but I don't think the market will move to this way doing things before the current AI wave has crested. So it is probably not good business sense to target my use cases. I have quite a lot of concurrency so I think my ideal hardware is a whole lot of little CPU cores with decent cache and Half Matrix Multiply Accumulate (HMMA) instructions. Once my CPU bottleneck becomes too painful I'll look at what options are on the market. I think, but am not sure, the early Tesla AI chips focused too much on optimizing ResNet and then later they went with CPUs with the HMMAs. Instead of using a variety of instructions I can instead approximate most functions with more HMMAs.
EDIT: changed my reference of matmul intrinsics to the more precise HMMA instruction
Very interesting point.
Back in 2015 I thought this would be the dominant model in 2022. I thought that the AI startups challenging Nvidia would be about that. Instead, they all targetted inference instead of programmability. I thought that a Tenstorrent hardware would be about what you are talking about - lots of tiny cores, local memory, message passing between them, AI/matmult intrinsics.
I've been hyped about Tenstorrent for a long time, but now that it is finally coming out with something, I can see that the Grayskulls are very overpriced. And if you look at the docs for their low-level kernel programming, you will see that Tensix cores can only have four registers, have no register spilling, and also don't support function calls. What would one be able to program with that?
It would have been interesting had the Grayskull cards been released in 2018. But in 2024 I have no idea what the company wants to do with them. It's over five years behind what I was expecting.
My expectations for how the AI hardware wave would unfold were fit for another world entirely. If this is the best the challengers can do, the most we can hope for is that they depress Nvidia's margins somewhat so we can buy its cards cheaper in the future. As we go towards the Singularity, I've gone from expecting revolutionary new hardware from AI startups to hoping Nvidia can keep making GPUs faster and more programmable.
Ironically, that latter thing is one trend that I missed, and going from Maxwell cards to the last generation, the GPUs have gained a lot in terms of how general purpose they are. The range of domains they can be used for is definitely going up as time goes on. I thought that AI chips would be necessary for this, and that GPUs would remain as toys, but it has been the other way around.
My customers won't touch non-commodity hardware as they see it as a potential vector for vendors to screw them over, and they're not wrong about that. In a post apocalyptic they could just pull a graphics card out of a gaming computer to get things working again which gives them a strong feeling of security. Having very capable GPU cards as a commodity means I can re-use the same ops for my training and inference which roughly halves my workload.
My approach to hardware companies is that I'll believe it when I see it, I'll wait until something is publically available that I can buy off the shelf before looking too closely at it's architecture. NVidia with their Tensor Cores got so good so quickly that I never really looked too closely at alternatives. I'm kind of hopeful that AMD SoC would provide a good edge compute option so I might give that a go.
I had a look at tenstorrent given this article and the Grendel architecture seems interesting.
I find this a very questionable business decision.
I am a big fan of Jim Keller, but this is semantics that argues CUDA the language is not a moat, when most people refer to CUDA the libraries.
CUDA has first party support in most of the libraries in use today, if you were to stray from that happy path you must have deep pocket to work through the hurdles of working with something like XLA.
There's a reason why Microsoft, Google, Amazon and Meta are still buying Nvidia accelerators even when they have their own in-house accelerators.
Server GPUs are the ones difficult to buy.
People don't really learn CUDA, they use it mostly through another library like PyTorch.
Try to make ROCm work on a random AMD card.
A good number of years ago I wrote my own Torch-like C++ NN framework using CUDA (cuDNN, cuBLAS) for the GPU, and one of the annoyances is that cuDNN has incomplete coverage of even the basic operators needed. Add/Sub/Min/Max/Sqrt/Negate are all provided as part of cuDNN, but if you want other common NN building blocks like Div/Exp/Log/Pow/Inv/InvSqrt then you have to write them yourself in CUDA and either forgo cuDNNs tensor-descriptor layout flexibility or re-implement that yourself too.
Of course frameworks like PyTorch support all the operators you'd expect, since they've written their own kernels in CUDA where the functionality is missing from cuDNN.
It's not clear exactly what Keller is referring to there. When people say CUDA they might be referring to the entire ecosystem (nvcc CUDA C/C++ compiler, CUDA API's, higher level cuBLAS, cuDNN, etc), or maybe just the base compiler (which lets you write your own kernels) and API for allocating memory, queueing kernels, etc.
The cuDNN kernels (convolution, etc) are highly optimized, as is cuBLAS (e.g. matmul), and I doubt anyone is going to do better writing these themselves. Does Keller consider using cuDNN as "writing CUDA" ?
As far as I'm aware the higher level, performant, CUDA libraries, as well as specialized components like TensorRT are written in CUDA, although that could mean a combination of C/C++ & PTX pseudo-assembler (ptxas is really a compiler, not an assembler). The alternative would be they they were written in hand optimized SASS assembler which afaik is only available outside of NVDIA via the Open Source CuAssembler.
I believe Mojo support for NVIDA is based on PTX. I'm not sure if that would be really be considered as "CUDA" or not if they are not using nvcc at all.
Most people, outside of framework vendors, would have no reason to use CUDA anyway, since it's just too low level. The only sane use case would be where writing a custom kernel in (e.g.) PyTorch or Mojo doesn't get the performance you want and you write than one kernel in CUDA. The hope would be that the Mojo compiler is good enough that this would not be necessary.
Can't help but think of that old yogi berra quote:
"Nobody goes there anymore, it's too crowded"
(but yeah, I suspect a lot of people use cuda via accelerated libraries/apps like opencv)
With CUDA you get a lot of flexibility which you don't get with Grayskull, at the expense of having a more layered/patched API which has been evolving for many years while R&D came up with new solutions to new problems.
I wonder what Nvidia's capabilities are in creating new, highly optimized cards, similar to Grayskull, but I wouldn't be surprised if they don't have any interest in creating them since their current products are already consuming all their resources.
It's great to see that companies like Tenstorrent are offering more optimized and pricewise more accessible products. It would also ease the situation with how hard it is to get Nvidia cards, because even companies doing only inferencing are buying chips which are developed to be able to do much more than that.
My understanding is that grayskull is a devboard for inference, Wormhole is supposed to do both, because you can scale it out with multiple cards.
> With CUDA you get a lot of flexibility which you don't get with Grayskull
Do you? From what I can tel tt-metalium gives you quite low level access.
Grayskull/Wormhole/... basically have a bunch of tensix cores, which each have 5 small rv32 cores, where some are used for data movement and some for computation. IIRC two (could be one) drive a SIMT compute unit.
From what I can tell there is no "matrix in matrix out" style fixed accelerator, the SIMT unit looks decently flexible, the grayskull one is a bit limited isa wise, wormhole is supposed to improve on that.
tt-buda is supposed to be the high level API with pytorch, tensorflow, ... support, but idk how much compatibility there is at the moment. I assume there is a lot of software work left.
Yet you got me to reading the following:
> t-Series Workstations
> Our t-series workstations are turnkey solutions for running training and inference on our processors, from a single-user desktop workstation in the t1000 up to the t7000 rack-mounted system designed specifically to function as a host with our Galaxy 4U Server
> 8 Grayskull Cards
> 16 Grayskull Chips
So these cards can also do training.
These are roughly analogous to compile time vs runtime for compiled programming languages.
Training is in general a more intensive task. However, in an ideal scenario training is run once and inference is run millions of times, so the lifetime cost of inference is bigger - this is why it might make sense to optimize for intense.
My printed copy (version 1.4) is from 2001, and that's a later version, I definitely had earlier, much slimmer books.
The 1.0 PDF reference on Adobe's web-site has a Copyright of 1993 and says "First Printing, June 1993".
https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandard...
And of course there were multiple unlicensed implementations. Famously, Apple switched from Display Postscript in Rhapsody and Mac OS X Server to Quartz and PDF in Mac OS X to remove the DPS licensing costs.
Editing the PDF just seems like the wrong place to place the edit since it is a format intended for interchange and archivation.
Layering changes on top makes total sense. Changing existing objects seems fragile. But maybe I feel that way because the format is known for being so complicated and bespoke that only Adobe's tools can edit it reliably. Just about any document format which can be converted to PDF is easier to edit, therefore it would make more sense to apply changes before conversion to PDF, if at all possible.
I agree that reading and scraping PDFs on scale is a very common operation.
I'm not an expert, but I imagine that they were created with Acrobat and it embeds some (proprietary) stuff in the PDF that only works with Acrobat. Acrobat is also the only PDF client that I'm aware of that supports signing PDF forms with certificates. There's other features in some PDFs that I've encountered that'll only work in Acrobat as well, like buttons embedded in the PDF to email the form or dropdowns in the form that don't work in other viewers.
It seems pretty swampy to me when 90% of the PDF forms I deal with can only be (fully) completed in Acrobat.
https://jasomill.at/HelloGoodbye.pdf
contains JavaScript that, per the spec, should display a pop-up alert ("Hello Acrobat <version>") when it is opened and another ("Goodbye Acrobat <version>") when it is closed.Actual behavior is inconsistent. On my Mac:
Acrobat: both alerts display
Preview, Safari: neither alert displays
Chrome, Edge, Firefox: only open alert displays
Moat as its typically used is to say that a business is protected by the moat, the prize being the castle/business. The residents/business are free to build and prosper, be productive in other words.
When Keller calls it a swamp, he's trying to say there's no castle to be productive in. The swamp makes it a mess, and thus hard for the business to actually maintain productivity. It is still difficult for the invaders to break in (?) but they don't need to...because it's a swamp. They should be busy making a better castle -- which is what Keller is trying to do.
Disagree with him or not, but I think that's the analogy.
Quite a thing to say as head of an AI hardware startup.
Most of the machine learning libraries we use/test in production (and that I use personally on my desktop) have hand written, extensively optimized CUDA kernels. It would be problem #1 switching to Tenstorrent hardware.
Not that I like that one bit, or that I really understand the optimization of these libraries. I do see triton code as well, but it still seems to be the suboptimal path.
I doubt he's saying that in the context of CFD or other similar ones.
Triton does seem way more approachable as something I'd like to pick up.
One advantage of CUDA, that Khronos/AMD/Intel realized too late, was the polyglot support that NVidia started to push and support since CUDA 3.0.
If CUDA is so good, it'd be great to know what it does that AMD cards can't. I've never gotten that far because I hit what seems to be some sort of kernel panic or driver lockup.
You are comparing Nvidia software with AMD hardware. It's the AMD software that has been lacking historically.
EDIT: here is a discussion from 2 months ago on CUDA versus ROCm. https://news.ycombinator.com/item?id=38700060
I've seen a mix of both for years. For the last 6 years the story has usually been some combination of "X program technically supports it, but you'll get 20% of the performance you'd expect given the hardware and/or you'll need to go through a bunch of hoops, and even then you'll have to troubleshoot tons of random errors and/or it'll crash the program or system"
That's on top of the "nobody supports ROCm, or if they do it's because of a single person who amd the PR - you're on your own though because none of the core contributers have amd hardware" which I'll admit is a chick-egg problem.
But given these two factors, it's always meant "if you want buy hardware to do data science, you have to go nvidia if you don't want to write the support yourself, and/or an insane headache that often resulting in switching to nvidia hardware anyway"
Gave up, bought an Nvidia card, zero problems.
I've seen something like this:
- Delphi desktop app (some time in the 2000s)
- Adobe Flash web app (in late 2000s to mid 2010s)
- 2015-era JS framework web app (mid 2010s to present, considered legacy codebase now, no feature development)
- 2020-era JS framework web app (late 2010s to present)
I expect another major rewrite or move to a new project in the next ~5 years
Is there anything 'inherit' to the instruction set that cannot be done anywhere else? I know that many Intel/AMD systems now have a bazillion cores, which is very handy for some things. Also, AMD seems to have a large numbers PCIe lanes, which is great for I/O between (e.g.) the network and on-system stuff.
x64 is here to stay with its incredible ecosystem and OSS friendliness - Coreboot, Intel and AMD GPU drivers, great choice of operating systems and apps with a fairly long usable lifespan even with Windows updates.
> currently there are actually no modern x86 CPUs on the market. Both Intel and AMD don't actually use x86 cores, but instead proprietary RISC cores, with microcode that translates the x86 code to RISC code on the fly at execution time.
https://cs.stackexchange.com/questions/132211/what-are-the-a... This cs.stackexchange link is a good read.
Wikipedia also states something similar,
> In the P6 and later microarchitectures, x86 instructions are internally converted into simpler RISC-style micro-operations that are specific to a particular processor and stepping level
I wonder if there's any ISA that could be written against the x86 cores that would be more efficient. In theory the chip could use that as a mode, so the chip itself wouldn't require an entire OS to shift before things could use it. I don't know anywhere near enough about the inside of an x86 core to have even a clue if such a thing would be possible. But it would be an interesting escape hatch from x86. Anyone who can flesh this idea out with knowledge of the x86 core internals is welcome to explain to me why my idea is bad and I should feel bad.
This feels like one of the stranger "no true Scotsman" arguments I've ever run across.
Even the 8086 had microcode that translated the instructions as generated by the compiler/programmer into the instructions that would be processed: http://www.righto.com/2022/11/how-8086-processors-microcode-...
I would love to know what the person who wrote that stackexchange answer would say in response to the question, "so which x86 processors used 'real' x86 cores?" Because from the very beginning of x86, there's been translation of the front-door opcodes into internal opcodes via microcode.
RISC is about the ISA, which is the interface between software and hardware. It is irrelevant whether the hardware is implemented with microops or hamsters running on wheels.
Yet if the existing microops used in popular CPUs made today were actually exposed as the interface between software and hardware, they would be VLIW ISAs, not RISC ISAs.
The whole "CISC have a RISC inside" idea is thus wrong at many levels, and just left over cultural damage from an Intel PR campaign ("risc vs cisc doesn't matter"), which was very effective at misinforming the tech world.
Note that Intel chips did no-doubt win in the market, but that was despite CISC, rather than thanks to CISC: Intel had a MASSIVE fab advantage, alongside the software moat around Microsoft OSs.
There is also the fact that RISC-ness is a property of the ISA, so applying it to describe a microarchitecture doesn't make much sense.
edit: reference: https://www.quora.com/Why-are-RISC-processors-considered-fas... (sorry for the quora link)
I used to say this all the time and I've been informed that it's something of a misunderstanding. For example, most RISC-V processors also decompose instructions into multiple μops: https://docs.boom-core.org/en/latest/sections/execution-stag...
So it isn't like there is a literal RISC processor inside the x86 processor with a tiny little compiler sitting in the middle. It's just that the out-of-order execution model requires instructions to be broken up into subtasks which can separately queue at the core's various execution units. Even pipelining instructions still wastes a lot of silicon (while you're doing an integer add the floating point ALU is just sitting there, bored) so breaking things up this way greatly improves parallelism. As I understand it, modern μop-based processor cores can actually have dozens of ALUs, multiple load/store units, virtual->physical address translation units, etc all working together asynchronously to chug through the incoming instructions.
This kind of factoid does more to obscure the truth than it does to illuminate it.
The truth of the matter is that all high-end CPUs do a µop translation, whether or not their frontend is a CISC or RISC ISA. Indeed, the very notion of CISC versus RISC is way overwrought in architecture textbooks, and this probably produces the garbled thinking: since everyone "knows" that CISC can't be superscalar, this means that the Pentium (in making superscalar x86) has to somehow be RISC.
Another thing to note is that there's not really anything called CISC. RISC is the overall term for a family of computer architecture design methodologies arising the 80's that argued for compiler-centric rather than assembler-centric design and simpler instructions, sometimes to the point that you omit hardware and call it a feature (e.g., delay slots). CISC is... everything else; it's a strawman constructed for RISC to compete against rather than a coherent design methodology.
In actual practice, though, RISC v CISC hasn't been relevant for decades. Some of the RISC design ideas have won out: there's generally a high emphasis on instructions that can be selected by the compiler over hand-tuned assembly, for example. But things like delay slots have been generally considered a failure. The architectures that are the most successful--x86 and ARM--are the ones that blur the line between RISC and CISC the most.
Actually, if you scrubbed the x86 assembly away and came up with some new assembly syntax (including new mnemonics of course), you could probably sell the x86 ISA as a "compressed RISC" ISA and get many people to believe you that it was designed as a RISC. x86 doesn't have many instructions that have crazy interrupt rules or multiple memory references (the string instructions are the main exceptions here), and it's this property which turns out to be really key to making something high-performance or not. It would be better for us to be honest about what enables or doesn't enable superscalar architectures rather than trying to argue that somehow x86 cheated its way to success.
Basically, I have noticed that people have been claiming RISC is the best for decades and yet real RISC architectures failed to beat x86's performance. Even worse, most of the RISC workstation and server makers exited the business, went bankrupt, or only sell systems to legacy customers who have not migrated to x86.
The other interpretation is that micro ops are actually a microarchitectural optimization and that RISC processors should use it too. The compressed instruction set of RISC-V is leaning towards that direction. RISC-V processors internally convert RISC-V instructions to RISC-V.
It is kind of hollow to talk about it, since it is hardly clear to say that this is a unique disadvantage or advantage.
I respectfully disagree. There's a major advantage of losing backwards compatibility with established ISAs, and that is freedom from licensing.
e.g. The main strength of RISC-V is not even its technical superiority; It is the free license.
The M1 and later big cores can dispatch four NEON FMA instructions per clock, so 512 bits worth of vector math, which compares OK with most Intel or AMD chips (Zen 4 can do two 256-bit MUL and two ADD, and Intel "client" bigcores since Sunny Cove typically do three 256-bit FMA).
If I have a core with four 128-bit neon vector units, I have the same throughput as an x86 with two AVX2 units or one AVX512 unit. However that 4x128-bit core is actually more flexible than the other two as I can do 4 different things at once, or 4 scalar operations per cycle. (Of course the downside is you spend more frontend resources on decode).
Given that most code isn't vector code, the multiple short vector length approach is actually superior on many common real-world workloads that aren't machine-learning (and CPU is unit-of-last-resort for large ML workloads anyway).
spoiler: it isn't
After the M1, the ARM companies needed to up their game, and Qualcom is rumoured to finally release something competitive to a real x64 chip in the near future. This isn't the first time they've made claims like these, though, and I very much doubt they'll live up to their promise.
Meanwhile, x64 has caught up to Apple in terms of compute power (especially per dollar, which is the reason x64 is so popular), is getting closer and closer to Apple's power consumption levels, and unless the M4 will have dramatically more performance, the ARM overtake will just have been an outlier.
I'm not so sure how long non-Apple ARM will be able to stay competitive given the turnaround x64 has managed to make in just a few short years. It looks like x64 still has plenty of room for improvement, and this proves that the reason compute performance has plateaued was that there was no real competition.
The difference now is that they actually have a set of customers that would buy a desktop-class chip if they produce one.
You ignore one huge class of x86 systems: servers. AWS Graviton launched 2 years before M1 and had competitive perf.
In engineering and software development, we have to accept trade-offs and imperfection. Every decision has pluses and minuses. Instead of name calling, Jim Keller should explain what his approach is and why he thinks it is an improvement. He should also explain its limitations and weaknesses. Finally, he should also explain how he will handle the inevitable compromises and imperfections which arise in successful systems.
RoCm = CUDA+cuBLAS+cuRAND
MiOpen = cuDNN
HIP is an AMD portability abstraction layer over the NVIDIA APIs that is pass-thru to cuDNN/cuBLAS/cuRAND on NVIDIA hardware, and pass-thru to AMD's equivalent APIs (that are basically drop-in replacements for NVIDIA's) on AMD hardware.
It is somewhat short-sighted though, since having support on all their GPUs would surely do wonders for adoption via more organic channels.
Cities are a swamp. The Web is a swamp. Linux is a swamp. Capitalism is a swamp. Democracy is a swamp.
Is it legal for China to hire Jim with an obscene pay package?
Why wouldn't they?
It's hard to imagine that this gaping hole was left wide open until the deft hand of Joseph Biden took the wheel.