Red Hat to author new Linux driver for Nvidia GPUs in Rust
phoronix.com
phoronix.com
I get that some of the architectural choices no longer make sense, and starting from scratch will address those. But is the goal to have performance that is somewhat comparable to the proprietary drivers? Or just good enough to run the desktop environment with hardware acceleration.
It is weird for a 3rd party to be maintaining a 2nd driver when the first party has a reasonable OSS driver.
- No telemetry.
- Enabling software blocked features.
- Emulation.
There are few possibilities about making a performant OSS driver:
- This is impracticable, because the software you want to run on the GPU is so complex, and the hardware is so complex, that you will never get enough insights to make something that compares with the people who can see it all.
- This is eminently practicable because: (1) the software is much simpler, and the GPU hardware much simpler. Perhaps there is a lot of obfuscation of the simplicity. Or (2) the application that best utilizes the hardware only needs a limited feature set that is within scope.
I'm leaning on "application limited scope" and "telemetry." It aligns best with what is actually happening, which is NVIDIA is scooping up a lot of valuable intelligence on LLM workloads; and that there isn't enough competition for LLM "ASICs" to make them cheap enough to be worthwhile.
Yeah, this is how, by writing your own driver. NVIDIA sells you turnkey DGX machines. It doesn't give you firmware. You have to be Internet-connected to refresh your various licenses, at some point, which is the moment the telemetry is shared. Google "NVIDIA telemetry."
If you are using NVIDIA on the cloud, well all bets are off. You are using their drivers. Amazon can't force you, in your VM, to install a different driver for the GPU you are using - there's no alternative to the proprietary one. Hence, my theory for why Red Hat could be paid to do this.
Telemetry is something that NVIDIA doesn't budge on for enterprises. You're welcome to see for yourself and start a sales call. I hear they're pretty busy.
This stuff is rather hush hush not to scare people but it's also well documented and certainly not an unfounded conspiracy. Its also why Europeans have adopted far reaching privacy laws (GDPR), their industries don't rely on consumer surveillance the way Americans have developed for the last decades.
Separately, most telemetry would tell stories like, "This project is a failure." Little incentive for people to adopt it. Then again, most OSS is ordered by the mania of its programmer-creators, not product or engineering quality informed by telemetry. Maybe in a Darwinian way, we only have the OSS that can thrive without telemetry and reactive product and engineering decisions.
Naturally the way they collect it is open source, it's largely de-identified, and because it open source you verify it's de-identified enough for you. And if it isn't you can turn it off.
So it's possible to do well, where "well" means gets you the data you without pissing off the users. Most proprietary don't bother to do it well for whatever reason.
But they should be careful: it a big factor in diving things like Home Assistant, Linux Desktop and now this, apparently.
Nvidia's driver cannot be included in the upstream linux as it doesn't follow kernel coding style and code organization, but more importantly it is tightly coupled to a single version of their GSP firmware - they have to be updated at the same time.
GPU drivers are significantly more complicated than any other drivers because they do things that are not considered part of a driver for any other device, the proprietary nV driver contains roughly as much code as the linux kernel.
You may be wondering why Red Hat is bothering with this effort then? I assume it’s so that the code can be added to Linux directly as opposed to being out of tree.
https://developer.nvidia.com/blog/nvidia-releases-open-sourc...
In 20 years when the current crop of greybeards are dead, how many experienced c programmers with kernel experience will exist.
Rust lowers the bar for contributing and also increases the pool of programmers who can.
So… what should we reckon? Would there a difference in getting new developers into `no_std` Rust in the kernel, and how different would that be versus having people learn freestanding C, with all the kernel add-ons and nicknacks?
I would still reckon that having familiarity with the standard rust (even with std) will still have more programmers willing to make the leap than learn C for this one project.
I know of less than a handful of C projects starting in 2024, I know there is a bunch in rust , even with no_std.
I had a lot of opinions like this too when I was 17, but age and experience has disabused me of many of them.
The embedded story on Rust still has many rough edges, but it's improving every day, and I could easily see it replace C in many places, given enough time (I wouldn't be too surprised if we eventually see companies distributing a BSP that's written in Rust).
Frankly, at this point in time, I think it's foolish to start a new project in C unless you have a really good reason to do so. Many embedded systems certainly qualify as a really good reason, but I very much hope that reason diminishes over time.
Some niches maybe but most packages are ancient and the cost of supporting rust Vs the gain is massive.
There are also way less people who can write rust. This forum with the "c is bad rust is the new God" attitude is not a representation of the world.
Sure, I didn't say I agree with them, I barely remember when I was 17, it was that long ago.
I'm saying that young people that I meet are more interested in c so the 'in 20 years no-one knows c' is not exactly true.
We try Rust now and then; it's not worth it yet in my opinion for what we do. The tooling and libs we have for c are vast and like said, c people are really easy to get, Rust not so much.
I hope this changes, but for now it's just too much of a struggle to warrant it. And I was only responding to the fear of not having capable c devs in 20 years. There will be plenty.
The Expressif folks are paying someone to make sure that Rust works well on the esp32.
Not sure I agree with that. C is a very easy language to learn. The problem with it is that you have to be careful with memory management. That does take effort to learn, and I still make mistakes after writing C for over 20 years.
I love Rust, but it is a very complex language, with a complex, full-featured stdlib, and rich, sometimes-inscrutable type system. It is more difficult to learn than C, but also more difficult (and sometimes impossible) to make many of the mistakes that you can make in C.
C looks easy on the surface, but the syntax is pretty dated and full of footguns (yes even just the syntax, not even talking about UBs), and learning the language is a pretty intimidating experience because every time you think you know something, you actually don't and get bitten later.
Rust on the other hand is a good language for CS students: you have a lot of things to assimilate upfront, but when you've reached the level required to fulfill the class, you're actually ready to use it in production, and the resulting code produced by a sophomore will be more stable than C code written by wizards.
To learn the language, yes. Definitely not an easy language to learn to write production software in.
Whereas an intermediate dev can easily get Rust basic in 2-3 months and can productively contribute to the safe portion of a complex project.
The ease of writing software in a given programming language is not a linear function of the complexity of the programming language.
Very simple languages are very difficult to use. Very complex languages are very difficult to use. Languages of intermediate complexity tend to be much easier to use than those at the extremes. C is more towards the "simple" extreme than the ideal, Rust is (IMO) a bit more towards the "complex" extreme than the ideal, but is closer to the ideal than C.
It reduces the amount of help that current experienced kernel developers can give to newer developers who want to write in rust.
The kernel could have been written in assembly with macros, it wouldn't make its development any harder. The best Rust can do is not make kernel more difficult than it already is.
The best Rust can do is enforce a number of invariants which people extremely experienced in writing "trivial C" still miss every day.
Rust is technically an improvement over memory unsafe language but it also has created enthusiam among largely young and proficient coders. The ecosystem has a lot of dynamism. Drawing those people in is one of the, if not the single most important thing for the health of projects going forward with how many leaders are close to retirement.
I can't talk about 'telco grade c', because I have never experienced it. I have however seen telco submitted code to upstream kernel and its not above the average quality.
> I wish I could be more optimistic about mutter!3304 coming soon, though... it's been 5 months already, and it doesn't seem anywhere closer to being merged into main To make matters worse, patched mutter 45.5 seems to be causing use-after-free on NVIDIA's kernel driver.
Christ could they just make Wayland usable first?
Unfortunately Linux Desktop still seems to be too irrelevant for NVIDIA. All current drivers have issues with suspend, multiple displays and Wayland among other things.
So Linux should be a priority for them, just not on the desktop. But the step should be small.
Now, for newer hardware, Nvidia has changed some aspects of the firmware and allows redistribution. So it's feasible to make a good open source driver.
AMD and Intel also use different drivers for different hardware generations, since eventually things change so much that it's better to start clean.
With regards to reverse engineering, Mesa has a number of reverse engineered drivers. That isn't anything new.
so nouveau gave up on it. they also expected nvidia to drop some firmware.
now that newer cards have GSP.bin firmware, which can be interfaced with easily - things are different. i would wager a guess that it's similar to atombios from amd. you just call a function in GSP and it knows what registers to poke with the right values to achieve what you need.
I wish I'd had more conversational courage at the time to "stand up" for nouveau with them; it was super rude to shit on someone else's hard work, especially when the reason why they had to do that hard work was because nvidia was a bunch of jerks with a shitty FOSS policy.
Presumably RedHat are in the same boat.
www.collabora.com/news-and-blog/news-and-events/introducing-nvk.html
NVK generates the commands to send to the hardware, by converting vulkan APIs to Nvidia-specific instructions, and feeds them to the kernel driver.
Tight control of software and licensing is a core part of their business model.
Not supporting most nvidia hardware is not a good thing. I hope it's adoption by commercial users (home users don't have the latest nvidia gpu generally) doesn't mean Nouveau bitrots from lack of companies with deep pockets paying for dev work. As for the inherent memory safety, well, remember how safe memory safe java was?
Memory-safe, system languages (eg Ada, Rust) address this by having as much as possible in memory safety with either little or zero runtime. Rust makes things safe that weren’t before. It has unsafe which is less used than some other languages. The type checker can also catch many logic-level errors. The safety situation is better with Rust than Java code.
That doesn’t mean it will solve NVIDIA users’ problems, though. They are usually worried about compatibility, performance, and reliability. Rust mostly helps in one area. We’ll see about the rest, especially compatibility, as you said.
I thought jetbrains used their own UI layer for their IDE, or am I confusing it with Eclipse there?
Swing feels pretty okay to me, at least in the times I've used it, especially when IDEs have GUI builders when you just want to do RAD.
I do wonder whatever happened to JavaFX, there was some hype around it years back but it doesn't seem like it got super widespread adoption: https://openjfx.io/
JavaFX was born as a scripting language[0], which most devs didn't like, and then Sun started the process to port it to Java.
In the middle of this, Sun went bankrupt and while Oracle took the JavaFX development further, they didn't see any value in adding yet another GUI framework to Java, and made it open source to the community by the time Java 11 was released.
A company, Gluon took over it [1], making their business case as means to use JavaFX to also target mobile OSes [3]. They also took over the JavaFX GUI designer. [2]
That was mostly ignored, as Java strong points focused on the server, Android is its own story, with their own frameworks, thus Swing was good enough for the market of desktop applications written in Java.
Additionally it didn't help that JavaFX is an additional dependency, with binary libraries, which kind of complicates the deployment story.
However, recently Oracle decided to be a bit more supportive of JavaFX, if it still matters, remains to be seen.
[0] - https://en.wikipedia.org/wiki/JavaFX_Script
[1] - https://gluonhq.com/products/javafx/
There were lots of old LCD monitors around the office. As soon as I could, I went dual screen. Then I decided to go triple-head.
Problem.
I found a 2nd card, another nVidia, but it was a different GPU generation. The same Nouveau driver can run it, but not the nVidia driver, and you can't have two nVidia drivers installed side-by-side.
I tried multiple nVidia cards from the office spares pile but I couldn't find two the same generation of GPU for ages. When I finally did, the next kernel upgrade nuked the nVidia legacy driver -- it only supports a certain range of kernel versions.
In the end, a colleague took pity and lent me one of his own cards, an old gamer's card with 4 outputs, 3 usable at once. Perfect.
But when he wanted it back, to sell it, it was back into driver hell.
In the end, I managed to get a huge fat double-slot AMD card from the IT department, with 4 outputs, and it Just Worked™ with a FOSS driver.
NVidia driver versions are a massive cluster-fsck and perfectly good working cards are now e-waste because nVidia doesn't maintain its proprietary drivers, doesn't support more than a handful of GPU families in each release, and won't let you install >1 driver at a time.
I am not a gamer. I don't give a rat's ass about 3D performance or CUDA or any of those toys. I just want a shedload of pixels in front of me, updated quickly, with some screens in portrait and some in landscape. (And a desktop that properly supports vertical panels so I can use those screens in any orientation I want with effective use of space.)
> Little hope of reclocking becoming available for GM20x, GP10x and GV100 as firmware now needs to be signed by NVIDIA to have the necessary access
1. "Nobody compiles a kernel these days". Sure, maybe not for daily use for most, but lots of people do for work and experimentation. Also, Gentoo.
2. "Everybody runs a modern CPU". Plenty of people still run decade-old hardware, and even older, because it works and in that decade objectives of PC have not significantly changed, so why would hardware?
On the other hand, the Rust language team should do their utmost to improve the speed of their compiler.
This driver is almost certainly not targeting 10 year old hardware.
So when you define your kernel config, you wouldn't include this driver. The new code would be entirely skipped over at compile time.
(Still think "rewrite everything in rust to solve all our problems" is a bit dumb.)
2010. I needed to do so for my Linux Kernel university class.
It took 8 hours on this piece of junk Pentium 4 I found in the school's e-waste bin. You know those Black Dell Optiplex machines that were everywhere like a plague.
They were days of such stress but also so much bliss.
Especially if you had some crap P4 that made modern gaming laptops look well-circulated.
AI has been flawless on Linux IMO(maybe not when it tries to use RAM instead of VRAM). When I use a server, I have 0 issues.
I used to get an occasional nvidia issue with 2 monitors and steam, but that went away at some-point.
It's a real shame that many things in Linux land are so badly designed/maintained that they have to be re-written from scratch every few years. The major exception being the base kernel itself I suppose.
amd uses the gbm properly, nvidia is half cooked and previously used eglstreams instead of because they did not like gbm. reality is every compositor just uses gbm.
So I guess the situation is not simply black & white but a bit more nuanced.
Nouveau supports seriously ancient Nvidia hardware. It's perfectly reasonable to make a clean break once in a while (both AMD and Intel have done it in their open source drivers).