AMD Open-Source GPU Kernel Driver Above 5M Lines, Entire Linux Kernel at 34.8M
phoronix.com
phoronix.com
Don't get me wrong I use the driver every day and AMD is definitely one of the good guys for making an open source driver and them who ported it are absolute heros. However.... Sometimes I wish AMD had tied down the isa to their cards a little better. Narrowed the interface if you would. because as it is the driver is so big because there is this combinatorial explosion of generated header files.
It may be more practical to rework the scripts to try to find ways to reduce the verbosity and redundancy. The actual .c driver code probably doesn't need every copy of every lines in all those .h files.
But I expect AMD would be skittish about open-sourcing anything that could even remotely be construed as HDL, even if it's just dry lists of registers. Open sourcing the drivers is one thing, but the hardware itself is another.
ROCm will be pointless until at least forward compatibility will be guaranteed by design.
C++ -> PTX -> Hardware ISA
So, presumably much of this is just lack of sufficient resources to assure that all the different combinations are optimally being compiled down on all the independent HW devices.
This was always the OpenCL problem too back before everyone just started using CUDA rather than trying to use it. It worked in a lot of places but the higher level code needed to be adjusted for individual devices because each vendor would do a reasonable job maintaining optimal code generation for their target given code optimized for their target, but switching vendors (AMD<->nvidia) would result in needing two different versions of the code tuned for the individual vendors ideas of how code should be tuned for their platform.
If that wasn't clear say you wrote and optimized code for vendor A's device X, when vendor A released device Y, things tended to work well, but moving that code to vendor B's devices basically required retuning/rewritting things.
> The competition is very much bound to a rather narrow ISA which is why CUDA is forward and backwards compatible whilst ROCm isn’t.
PTX isn't an ISA and it goes through another translation step before it hits the real ISA (SASS), which does change internally. Though you are right that PTX as intermediate is part of the reason for CUDA's success in that it makes the code portable across cards.
Take amd64 as a counter example, not only does it not take 4 million lines of code to interface with an amd64 processor, when something new is added the processor is still compatible with existing code. x86 as an architecture is a bit of a train wreak but much of it's success is the effort to keep the interface stable.
Also for example arm, this is getting better but the hardest part about using arm is the isa is all over the place, every single bloody SBC appears to have a different isa.
No one currently or in the decades past has done otherwise.
I raised an eyebrow but I have only the vaguest notion of how the hardware works and what a driver might have to manage.
Probably if AMD wanted to spend the time, they could compress it down to a fraction of it's current size.
0: https://github.com/spdk/spdk/blob/master/include/spdk/nvme_s...
Even on small m.2 style standalone drives, your looking at code, which not only handles the details of managing flash error correction, wear leveling, garbage collection, etc, etc, but all the code required to manage the thermal, voltage, pcie link training, etc of the 2-5 or so microcontrollers embedded in the drive and possibly an RTOS or two hosting it all.
Never mind fabric attach (DPU?) NVMe devices which do all that, plus deal with thin provisioning, partitioning, deduplication, device sharing, replication, RAID, etc, etc. Frequently themselves embedding a Linux (or similar level of complexity OS) kernel in the control plane.
GPU or CPU? If talking about the latter only two [four] should count (ARM & x86 [+ [* 2 64BitVersion]]. If you meant the former forget my comment.
I'm not familiar with AMD's GPUs, but have done some "bare metal" Intel GPU programming, and there's definitely a lot of commonality between different generations going all the way back to the i810.
ehh, no.
almost all of this are header files
>Meanwhile the open-source NVIDIA "Nouveau" driver is around 201k (21.7k blank lines, 24.3k lines of comments, and 155k lines of code). Or the Intel i915 DRM kernel graphics driver is around 381k lines via the same cloc judgment.
so it seems like GPU driver is around 1% of kernel's code
and you start thinking why actually kernel has this much code if GPU (out of all software) needs just around 1%.
Nouveau will bring up you're Linux desktop just fine and connect all your monitors.
It will run OpenGL apps and light games.
Last time I checked it could not boost the clock and push most chips to high performance mode.
This is mostly 100% Nvidia's fault of not opening up the specs.
Otherwise what nouveau has been able to reverse engineer is just amazing.
The nouveau driver is so much less painful to use, it just works completely seamlessly. If you don't need the extra performance, it is not worth the trouble to install the Nvidia blob.
The reason is, GPU drivers are basically complete operating systems, just for the secondary computer we call the GPU instead of the CPU.
The large majority of the "Linux kernel code" is drivers, but the large majority of the driver code is the GPU drivers. And the GPU doesn't have its own collection of drivers for every network chip or USB controller ever made.
The second biggest part of the Linux kernel is "arch/" which is the architecture-specific stuff, but GPUs don't really have that either -- a given vendor more or less corresponds to a platform architecture, but if you compare it to, say, "x86/" that's only ~11% of "arch/" and <1% of the kernel.
The reason the GPU drivers are so big is that they're gibberish. Instead of specifying an interface to interact with the GPU in a sane way, they're full of magic numbers that seem to map a (large, overly complex) API interface into the values you pass to the GPU to call the API functions implemented by the GPU's firmware, which is the actual GPU "operating system" but the vendors want to keep as a black box.
Which is obviously counter-productive because it keeps users from optimizing for their GPU, which would make things run faster on it (or have fewer stability bugs), which would make more people want to buy them over a competitor's, which was supposed to be the reason for the secrecy.
If true, AMD is doing something wrong here. And yes, giga tons of generated headers related to registers.
Most of the GPU software stack can reasonably be outside the kernel. There's no obvious reason why much of it would need to be in ring-0 where bugs cause OS crashes and security vulnerabilities.
> Meanwhile the open-source NVIDIA "Nouveau" driver is around 201k (21.7k blank lines, 24.3k lines of comments, and 155k lines of code). Or the Intel i915 DRM kernel graphics driver is around 381k lines via the same cloc judgment.
But without separate line counts of the generated data tables and actual human written code, we don't have the real line count numbers to compare.
FreeBSD: ~9M loc
NetBSD: ~7M loc
OpenBSD: ~3M loc
And this includes the base userland (not just kernel)https://www.csoonline.com/article/564373/is-the-bsd-os-dying...
Maybe the figures you quote exclude things imported from elsewhere like gcc and llvm, I get a figure of 75M loc for base + kernel of NetBSD-10.
So what's the point of saying that it's large?
Pointing out that enormity is important because source files need to be stored; interpreted, versioned and parsed by humans/IDE's. It has an externalised cost (but, then again, isn't capitalism all about externalising costs?)
AMD maintains it but do we know how they are generated? Probably not.
It's like a gift that stinks but you can't complain about because it's a gift.
Having worked in the semi industry, I can fathom a guess: It's a spaghetti mess of cascading Perl scripts that parse the Verilog/VHDL design files, with their development going back 20+ years, full of comments like "don't touch this line because it breaks another line, nobody knows why", and maintained by a team where a gray-beard "Gandalf" engineer wearing an ATI t-shirt, has most of the deep-down low-level knowledge on how to un-fuck them whenever they get fucked, pardon my french.
Curious how much of AMD Radeon GPU development now is being done in Markham-Canada, as AFAIK, the modern Radeon architecture stems from ATI's acquisition of ArtX[1], a US-based spin-ff of SGI, which was responsible for the GPUs in the Nintendo GameCube, Wii and many other innovations like programable shaders, later found in ATI/AMD GPUs.
>Much of that team is European now though from what I hear, and they’re generally a good bunch.
I didn't know AMD has a GPU design team in Europe. Where? I know they had a fab in Germany and they have an office for the Ryzen and Infinity Fabric R&D in Romania, but I had no idea they do GPU stuff as well in Europe. Where is that office?
AFAIK they don't, but the Linux driver guys seem to be mostly German and Polish and such. And yeah, they are doing good work. I half-expect AMD to reboot their Windows driver from the Linux driver code base at some point.
Basically those files are generated from AMD GPU register data files where majorify of registers are documented, but there of course bunch of magic numbers as well probably because they belong to HDCP or other cases where documentation only available under NDA.
There been a number of leaks of AMD internal documentation so anyone who is into GPU drivers can really find a lot of information on their GPU internal workings.
I've archieved some of it many years ago and it's was never DMCA*ed:
https://github.com/ArseniyShestakov/rai-bonaire
Source was a talk on CCC.
Whereas I worked on Chrome's V8 C++ code for a year and I still could not say I understand more than half of it. Its complexity is a factor more than the Linux kernel.
If they were generated as part of the build, they would not be counted as SLOC (not being "source").
Should someone decided that they'd start working through the code, removing duplicate code and clean up headers, functions and abstraction, they work would either be rejected, or undone with the next AMD code dump?
Sure, you can fork the code if you really feel that strongly about it. My main "issue" is that it basically removes one of the big benefits of open source, that we can collaborate and do better as a collective. If it's just a big code dump that other kernel developers can't really touch it's more "source code is available" than actual open source.
When you send someone a PR, you are demanding that they do work for you to review and merge. Open source licensing does not mandate that they do this. Heck, most open source licenses even disclaim warranty to avoid obligating the authors of even doing work that the law would otherwise require them to do. Now yes, some people will help you with problems. This is because they're nice, not because it's open source.
Because it is open source, but not open contribution.
Open source vs. open contribution are orthogonal concepts: the former is about licensing, the latter about the organization of the development process.
Wouldn’t it be better to load them in such a way that a crash in the GPU driver can be recovered from as opposed to crashing the whole system?
Other operating systems load the GPUs drivers separately.
Depending on your definition of 'the whole system' this is not entirely true.
The AMD driver for example will reset the GPU and force a xorg restart: https://paste.debian.net/plain/1290344
Now this does mean all desktop applications don't close properly, so I restart the PC to a 100% sure stability is at 100%, it didn't cause a kernel panic like the GPU crash would do previously.
I've had this crash happen only twice in ~90 hours of play time.
GPU: 5700XT
Driver: Mesa 23.1.5
And there are great tools for exploring directories of files, my current favorite is dolphin with two or three panes.
[1] https://upload.wikimedia.org/wikipedia/commons/d/d5/GNOME_Di...
[1] https://en.wikipedia.org/wiki/Linux_kernel#/media/File:Sanke...