Giving Rust a chance for in-kernel codecs
lwn.net
lwn.net
And yes, you can parse safely in C. Unfortunately, correctness in C has been proven to be rather difficult to achieve.
The other main tradeoff is that explicit lifetime annotations make (big) refactorings or explorative coding harder, unless one uses a lot copies or shared pointers (unacceptable for Kernel code).
However, I am not familiar how those two tradeoffs ought to be solved in Kernel. I would expect that Rust would do rarely changing high level abstractions expected to never change fundamentally in design and/or being limited in scope.
SOME systems have write protection logic on the RAM controllers themselves, but this is not universal.
How can one tell if a system has RAM controller based security, what name does this write protection go by?
Ubuntu 22.04 tried to turn it on by default but switched it off again due to random mostly graphics related regressions on random hardware: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1971699
Some other history: https://www.phoronix.com/news/Intel-IOMMU-Gfx-Default-Try
We really do need it though. I am always reminded of the very old Apple "Firewire Memory Bypass" which rendered flames to the screen just by plugging a firewire device in - because firewire had direct and originally unprotected DMA access: https://www.pentestpartners.com/security-blog/hack-demo-vide...
It is for this reason that even without IOMMU, as a workaround, you have to often give permission to thunderbolt devices to connect. Some details on that here: https://wiki.archlinux.org/title/Thunderbolt
There is also a small but noticable performance hit to using the IOMMU, not so noticable on a general setup but if you are doing high-speed disk & network I/O like ceph storage in excess of 10Gbit/s or millions of IOPS you will notice it. You can Google that.
You can also run into other weird behaviour, for example when using kdump to create a kernel crash dump it will kexec from the old kernel into a new kernel to produce the crash dump. The system doesn't go through a firmwire/uefi/bios reset so the hardware state of network cards, etc, doesn't get reset. So if you have any hardware driver state that isn't properly reset, you might for example have your network card DMA a packet directly into host memory in the time window before it gets reset. With IOMMU that might trigger errors, with it off it will hopefully not overwrite anything important but may also overwrite something important :)
These things are all of course fixable, but since it's still off by default much of the time, lots of these bugs persist for a long time.
Disclaimer: I am not an expert in this area it's just anecodtes from my life as a Linux Geek & Support Engineer. Should be about 90% accurate but I am sure I glossed over some solid details :)
At any rate, it's unfortunate that entire media file formats have to run in kernel space in order to implement hardware acceleration. There's no better way to do it?
Sorry. I hate it too.
It is analogous to passing a pointer to a buffer to a library function so that the library can write the contents of the buffer directly. You rely on the library (device) operating properly and not writing outside of the designated buffer. That is the baseline.
As you state, you could add a enforcement mechanism that defines the extents of memory the external actor is allowed to store like how many language runtimes check and disallow out-of-bounds accesses. However, if you have multiple outstanding buffers, control structures, or other complex device-accessible structures then enforcing precise "bounds" checking rapidly demands very complex "bounds" definition/enforcement. Language runtimes can do this because they support arbitrary code, but enforcing in hardware that, say, every access lies within the nodes of a "user-constructed" red-black tree rapidly becomes infeasible.
You basically get one of two options at that point, either you rely on the hardware working properly and do nothing, or you design your driver to only require coarser isolation that fits within hardware-definable boundaries. Most opt for the former. If you do the latter then there are various ways of actually enforcing the isolation such as IOMMUs or you could have a DMA controller that basically functions as a IO MPU (define ranges the device can access) (I am not actually aware of any DMA controllers that actually do this as a security measure, but it is theoretically possible).
You do have to be careful that they actually enforce the isolation. For instance, I question if the x86-64 IOMMU implementation is actually safe against a malicious device due to certain supported features such as device IO TLBs, but I do not know about the actual hardware implementation to know for certain.
That is: why not sandboxing rather than a rewrite?
Here's a blog post that demonstrates embedding an entire Go program as a blob into a kernel module and running it in userspace from the module:
Memory safety can be addressed via kernel tools or frameworks and should not be the job of a language IMHO.
And Rust does seem to be increasing the rate and level of support small outfits can offer - see Asahi Linux for an example of that.
if your usage requires that fine. don’t place your requirements on other users.
also you are basically saying adopting linux is only for big tech. i don’t think that is in the spirit of open source.
> if your usage requires that fine. don’t place your requirements on other users.
Conversely, if you don't care about code quality feel free to not use Linux or make your own in-house patches without contaminating the kernel other people are trying to use.
> also you are basically saying adopting linux is only for big tech. i don’t think that is in the spirit of open source.
If only "big tech" can write kernel code safely then so be it. Of course, I think that premise is nonsense, but I don't care about the onramp enough to compromise on security.
Have you written much C or C++? What kind of kernel tools or frameworks are you thinking of? Have you ever used Rust? Are you familiar with the different kinds of memory errors?
I really struggle to imagine how anyone who is actually familiar with all this stuff could say things like this but you aren't the first...
i don’t need or want rust.
I think most C++ people will (after using it for a while) begrudgingly admit that it is a very nice language, at least in some regards.
So do you just not mind the frustration and time wasted debugging these issues?
Memory safety could've been addressed through kernel tools and frameworks for decades, but it hasn't. And it _should_ be part of the language, as even low level languages like C try not to clobber memory and specify undefined behaviour in cases where you may accidentally end up doing it anyway.
There are good arguments for and against other Rust features such as the way panic!() works and the strictness of the borrow checker. However, "C and tooling can do everything your fancy pants new language does" has been said for longer than I've been alive and yet every month I see CVE reports about major projects caused by bugs that would never have passed the Rust compiler.
Out of every reason I can think of, a lack of hardware support seems like the least likely reason for a hardware manufacturer not to upstream hardware support. Look at companies like Qualcom, with massive ranges of devices and working kernel drivers, hacked together because upstreaming doesn't benefit them. Look at companies like Apple, who doesn't care if their software works on Linux or not. The Linux kernel supports everything from 90s supercomputers to drones to smart toothbrushes, there's no lack of hardware support.
it’s clear from your comment you never built a device yourself running linux on an obscure SoC dealing with patches that were never accepted.
Applying patches to the Linux kernel isn't that hard? Porting abandoned code sucks, but putting in the work to finish and include the patches would just move that annoying work to the rest of the kernel instead of a few people working on weird SoCs.
> essentially you are saying rust is great because companies can hire clueless people and rust will compensate for their inability
That's not what I'm saying, at all. What I'm saying is that nobody is perfect, not even C developers. Even if the mythical perfect C developer, who writes bugless code, does all the necessary sanitization, and always applies the necessary layers of tooling even when personally inconvenient, does exist, they can make better use of their time by using a language toolset that doesn't necessitate that extra work in the first place, regardless of whether that's Rust, C++, Zig, Frama-C, Carbon or Jakt.
The "everybody who makes mistakes is clueless" crowd is exactly why C developers get such a bad rep. C compilers have a billion warnings and -Weverything -Wnoimeaneverything -Wnoeverysinglething exactly because C developers, like any other developers, need help.
I know why people choose Rust today, but we could have had all of those benefits decades ago if not for the recalcitrance of reflexive C++ haters.
That said, Linux is a very conservative and idiosyncratic project. It's really the bazaar. (No issue tracker, firehose of emails, etc.)
They're arguably still not there yet, sure the standard was ratified years ago but the implementations are still a mess and uptake is almost nonexistent.
For instance most of the C++ stdlib is useless for kernel (or embedded) development, so you end up with a C++ that's not much more than a "C with namespaces". At least Rust brings a couple of actual language improvements to the table (not a big fan of Rust, but if the only other option is C++, then Rust is the clear winner).
Zircon is an existence proof that you can use the C++ std library in a kernel. Do they use every damned thing in std? No, but they do use array, pair, unique_ptr, iterators, and more.
Except the bits allowed in a freestanding environment and (stretching the definition of "used in") the "nolibc" partial C standard library the kernel includes for environments that don't have any other libc available (`tools/include/nolibc`, mostly used for kernel tests where there's no userspace at all).
I find the C++ equivalent of the Rust iterators and such to be even harder to read (almost an accomplishment, given the density of Rust code); I don't think features like ranges would be expressed more succinctly using std::range the same way it can be done in Rust, for instance. I also find C++'s iterators' API design rather verbose, and I don't think there's much good to be said of implementing generics through C++ templates. There are good reasons for why the language was designed this way, but I get the impression succinctness didn't seem to be a primary objective designing them. Rather, I get the feeling that the language designers prioritised making the features accessible to people who already knew C++ and were used to the complexer side of C++.
I don't reflexively hate C++, just all the implementations of it.
Once you refuse their fantasy "Don't lift a finger" way to add C++, they lose interest.