The Dirty Pipe Vulnerability
dirtypipe.cm4all.com
dirtypipe.cm4all.com
> Blaming the Linux kernel (i.e. somebody else’s code) for data corruption must be the last resort.
^^^ I can only image the stress levels at this point.
Circa 10 lines of C. Beautiful.
First it helps to triage whether the bug is in your own code or in other code.
You also then can use it like the author and find the commit which introduced the bug.
And you can use it as a test case to verify if the bug is closed.
wish more would be written like this
Every few months, someone would always come up with a reason why "bad entropy" was the cause of a bug.
It never was, but it felt like a thing we had to go through, to get to the real cause.
I was working in MSFT back than and I was writing a tool that produced 10 GB of data in TSV format , that I wanted to stream into gzip so that later this file would be sent over the network. When the other side received the file they would gunzip it successfully, but inside there would be mostly correct TSV data with some chunks of random binary garbage. Turned out that pipe operator was somehow causing this.
As a responsible citizen I tried to report it to the right team and there I ran into problems. Apparently no one in Windows wants to deal with bugs. IIt was ridiculously hard to report this while being an employee, I can't imagine anyone being able to report similar bugs from outside. And even though I reported that bug I saw no activity in it when I was leaving the company.
However I just tried to quickly reproduce it on Windows 10 and it wouldn't reproduce. Maybe I forgot some details of that bug or maybe indeed they fixed this by now.
There are lots of things which have been fixed in Windows 10, I'd go so far to say 1903 (19H1) is where things started to settle down, but even the latest versions are not perfect. When the Israeli/Palestinian conflict broke out in 2019, some of the US military computers started playing up for about a week, after the US vetoed something at the UN level regarding this conflict. So MS still has a long way to go to get things secure.
Glad it looks like they got around to it though.
This gives attackers an advantage (they are incentivized to read commits and can easily see the vuln) and defenders a huge disadvantage. Now I have to rush to patch whereas attackers have had this entire time to build their POCs and exploit systems.
End this ridiculous practice.
(This is not a rhetorical question. I can possibly influence this policy, but unsubstantiated objections won’t help.)
Evidence of this being helpful to attackers and not defenders? IDK, talk to anyone who does Linux kernel exploit development.
edit: There you go, Greg linked his policy, which explicitly notes this.
1. The commit message [1] does not mention any security implication. This is reasonable, because the patch is usually released to the public earlier and it makes sense to do some obfuscation, to deter patch-gappers. But note that this approach is not a controversy-free one.
2. But there is also no security announcement in stable release notes or any similar stuff. I don't know how to provide evidence of "something simply does not exist".
3. Check the timeline in the blog post. The bug being fixed in stable release (5.6.11 on 2022-02-23) marks the end of upstream's handling of this bug. Max then had to send the bug details to linux-distros list to kick off (another separate process) distro maintainers' response. If what you are maintaining is not a distro, good luck.
Is this wrong-sounding enough?
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
#2: upstream makes no general effort to identify security bugs as such. Obviously this one was known to be a security bug, but the general policy (see #1) is to avoid announcing it.
#3: In any embargo situation, if you’re not on the distribution list, you don’t get notified. This is unavoidable. oss-security nominally handles everyone else, but it’s very spotty.
Sometimes I wish there was a Linux kernel security advisory process, but this would need funding or a dedicated volunteer.
I guess a Linux kernel security advisory process is needed to fix this, but yeah :(
This is already happening https://osv.dev/list?q=Kernel&affected_only=true&page=1&ecos...
No, not at all, just to leave time to users to deploy the fix before everyone jumps on exploits. This is important because every single backported patch is a candidate for an exploit already, and it's only a matter of time before any of them is exploited. Reason why embargoes have to stay short. It takes some time to figure whether a bug may have security impacts. It takes much less time once this is figured, to develop an exploit.
By the way it could have really happened that the fix for data corruption would have been merged first, and only later the author figured there was a security impact. And the patch wouldn't have been any different. That's why leaving 1-2 weeks for the fix to flow via distros to users, and having the author post a complete article is by far the best solution for everyone.
Thus the only solution is to have a public fix describing the bug and not necessarily all the details, while distros prepare their update, and everyone discloses the trouble at the same time. Those who need early notification MUST ABSOLUTELY BE on linux-distros. There's no other way around. As soon as the patch is published, the risk is non-nul and a race is started between those who look for candidate fixes and those who have to distribute fixes to end users.
This is not about silently patching or hiding bugs, quite the opposite, it's about making them public as quickly as possible so that the fix can be picked, but without the unneeded elements that help vandals damage shared systems before these systems have a chance to be updated. Then it is useful that the reporter communicates about their finding, this often helps improve general security by documenting how certain classes of bugs turn to security issues (Max did an awesome job here, probably the best such bug report in the last few years). And distros need to publish more details as well in their advisories, so details are not "hidden", they're just delayed during tha embargo. Those who are not notified AND who do not follow stable are simply irresponsible. But I don't think there are that many doing that nowadays, possibly just a few hero admins in small companies trying to impress their boss with their collection of carefully selected patches (that render their machine even more vulnerable and tend to make them vocal when such issues happen).
In addition it's important to keep in mind that some bugs are discovered as being exploitable long after being fixed. That's why one MUST ABSOLUTELY NOT rely on the commit message alone to decide whether they are vulnerable or not, since it's quite common not to know upfront. I remember a year or two ago someone from Google's security team reported a bug on haproxy that could cause a crash in the HPACK decoder. That was extremely embarrassing as it could allow anyone to remotely crash haproxy. We had to release the fix indicating that the bug was critical and that the risk of crashing when facing bad data was real, without explaining how to crash it (since like a kernel it's a components many people tend to forget to upgrade). Then after the fix was merged, I was still discussing with the reporter and asked "do you think it could further be abused for RCE?". He said "let me check". A week later he came back saying "good news, I succeeded". No way to get that info in the commit message even if we wanted to, since that was too late. Yet the issue was important.
Speaking of Brad, I personally think that grsec ought to be on linux-distros, but maybe they prefer not to appear as "tainted" by early notifications, or maybe they're having some fun finding other issues themselves. We even proposed Brad to be on the security list, because he has the skills to help a lot and improve the security there. He could have interesting writeups for some of the bugs, and it would probably change his perception of what happens there. Maybe one day he'll accept (still keeping hope :-)).
If you wish to disagree with how we handle all of this, wonderful, we will be glad to discuss it on the mailing lists. Just don't try to rehash all the same old arguments again, as that's not going to work at all.
Also, this was fixed in a public kernel last week, what prevented you from updating your kernel already? Did you need more time to test the last release?
Edit: It was fixed in a public release 12 days ago.
> Just don't try to rehash all the same old arguments again, as that's not going to work at all.
No shit. People have been trying to explain all of this to you for decades lol I'm not stupid enough to think I'll succeed where they've failed.
It's been working semi-well, and gives us a way to deal with longer embargo times (like months instead of weeks and days), but it does not integrate well into the linux-distro-like way of working just yet, which is an issue that hopefully will be resolved sometime in the future if the linux-distro members wish it to be.
Here the goal was to make sure that all those who correctly do their job are fixed in time. And they were. Those who blatantly ignore fixes... there's nothing that can be done for them.
I've already explained that most people have some sort of cadence for updating.
Based on your other comments you're just going to ignorantly parrot Greg's talking points. I don't think you have much insight into this.
I don't think you know who wtarreau is.
> Linus summarized the reasoning behind this behavior in an email to the Linux Kernel mailing list in 2008 ...
Since severity can be a moving target, it seems like there is no straightforward solution. With that said, by hiding the known ones, older distros don't have much of a hope in hell of getting all reported CVE fixes back-ported.
Why isn't there a public index mapping known CVE fixes to git commit IDs? This seems totally doable and would make the world a more secure place overall.
Older distros have always had a ton of privilege escalation bugs and I don’t think that’s ever gonna change. If you can’t keep everything updated, your machines have to be single-tenant.
Attackers don't just crawl code at random. Starting with known crashes or obfuscated commits is always faster.
If you do not want to rush to patch more than you have to, use a LTS kernel and know that updates matter and should be applied asap regardless of the reason for the patch.
When someone submits a patch for a vulnerability label the commit with that information.
> You have to rush to patch in any case.
The difference is how much of a head start attackers have. Attackers are incentivized to read commits for obfuscated vulns - asking defenders to do that is just adding one more thing to our plates.
That's a huge difference.
> the logical step is that it doesn't require immediate action when the label is not there.
So I can go about my patch cycle as normal.
> Never mind that the bug might actually be exploitable but undiscovered by white hats.
OK? So? First of all, it's usually really obvious when a bug might be exploitable, or at least it would be if we didn't have commits obfuscating the details. Second, I'm not suggesting that you only apply security labeled patches.
See also: https://twitter.com/spendergrsec
Your last point is wrong. Simple example, which of the following thousand bugs are exploitable? https://syzkaller.appspot.com/upstream
If you can exploit them, you can earn 20,000 to 90,000 USD on https://google.github.io/kctf/vrp
But they clearly don’t want nor care about making that call (and even more clearly basically expect everyone to run the latest kernel at all times (and if you run into a bug there no doubt you’ll be told to not run the latest kernels).
But going about your normal patch cycle as normal for things not labelled "Security Patch", just means if the patch for some reason should have been tagged but wasn't, you're in the same situation.
I do see the value in your approach, but it just does not change anything for applications where security is top priority.
Well Xen for instance includes a reference to the relevant security advisory; either "This is XSA-nnn" or "This is part of XSA-nnn".
> If the maintainers start to label commits with "security patch" the logical step is that it doesn't require immediate action when the label is not there. Never mind that the bug might actually be exploitable but undiscovered by white hats. If you do not want to rush to patch more than you have to, use a LTS kernel and know that updates matter and should be applied asap regardless of the reason for the patch.
So reading between the lines, there are two general approaches one might take:
1. Take the most recent release, and then only security fixes; perhaps only security fixes which are relevant to you.
2. Take all backported fixes, regardless of whether they're relevant to you.
Both Xen and Linux actually recommend #2: when we issue a security advisory, we recommend people build from the most recent stable tip. That's the combination of patches which has actually gotten the most testing; using something else introduces the risk that there are subtle dependencies between the patches that hasn't been identified. Additionally, as you say, there's a risk that some bug has been fixed whose security implications have been missed.
Nonethess, that approach has its downsides. Every time you change anything, you risk breaking something. In Linux in particular, many patches are chosen for backport by a neural network, without any human intervention whatsoever. Several times I've updated a point release of Linux to discover that some backport actually broke some other feature I was using.
In Xen's case, we give downstreams the information to make the decisions themselves: If companies feel the risk of additional churn is higher than the risk of missing potential fixes, we give them the tools do to so. Linux more or less forces you to take the first approach.
Then again, Linux's development velocity is way higher; from a practical perspective it may not be possible to catch the security angle of enough commits; so forcing downstreams to update may be the only reasonable solution.
Currently the situation is that you can just follow development/stable trees right (e.g. [0])? Why would you only want the security patches (of which there look to be a lot just in the last couple weeks). Are you looking to not apply a patch because LK devs haven't marked it as a security patch?
[0]: https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux...
On the other hand, just on that page I linked, there's... a lot of issues in there I would consider patching for security reasons. I don't know how reasonable it is, given the existing kernel development model, to tag this stuff in the commit. The LTS branches pull in from a lot of other branches, so like, which ones do you follow? When Vijayanand Jitta patches a UAF bug in their tree, it might be hanging out on the internet for a while for hackers to see before it ever gets into a kernel tree you might consider merging from.
I guess what I'm saying here is that it seems like a lot to ask that if I find a bug, I:
- don't discuss it publicly in any way
- perform independent research to determine whether there are security implications
- if there are, ask everyone else to keep the fix secret until it lands in the release trees with a [SECURITY] tag
- accept all the blame if I'm ever wrong, even once
That too is a lot of overhead and responsibility. So I'm sympathetic to their argument of "honestly, you should just assume these are all security vulns".
So maybe this is just a perspective thing? Like, there are a lot of commits, they can't all be security issues right? Well of course they can be! This is C after all.
Like in that list, there's dozens of things I think should probably have a SECURITY tag. Over 14 days, let's just call that 2 patches a day. I'm not patching twice a day; it's hard for me to imagine anyone would, or would want to devote mental bandwidth to getting that down to a manageable rate ("I don't run that ethernet card", etc.)
So for me, I actually kind of like the weekly batching? It feels pragmatic and a pretty good balancing of kernel dev/sysadmin needs. Can I envision a system that gave end-users more information? Yeah definitely, but not one that wouldn't ask LK devs to do a lot more work. Which I guess is a drawn out way of saying "feel free to write your own OS" or "consider OpenBSD" or "get involved in Rust in the kernel" or "try to move safer/microkernel designs forward" :).
For this special case, yes. But for the vast majority of bugs it's the opposite and existing bugs get exploited later, thanks to some people who think that some patches are not security-related and do not apply the fixes.
To me, it seems like the average corporate security team is not going to worry about these kinds of attackers. Security for state secrets might, but they seem likely to be clued in early by Linux developers.
I'm probably missing something tho.
Virtually every single Linux user. I think what you're missing is how commonplace and straightforward it is for attackers to review these commits and how uncommon it is for someone to be on the receiving end of an embargo.
Most exploits are for N days, meaning that they're for vulnerabilities that have a patch out for them. Knowing that there's a patch is universally critical for all defenders.
For context, my company will be posting about a kernel (then) 0day one of our security researchers discovered. You can read other Linux kernel exploitation work we've done here: https://www.graplsecurity.com/blog
I get that every linux user could be attacked. But why would someone with the relevant knowledge that could pull this off attack a given linux user? Why are you worried about it? (Not trying to be sarcastic, trying to get a sense of what threats you are worried about).
If the patch was flagged as a security problem from the beginning, it would give advantage to attackers, since they would know that the particular patch is worth investigating, while the defenders would have to wait for the patch to be finalized and tested anyway.
Their point is that a full-time attacker (and there's enough money in it to do it as a full-time job these days) can look for obfuscated commits and take the time to deobfuscate them, whereas a defender doesn't have that kind of time.
My point was that if security patches are flagged as such from the start, it saves attackers lot of time (and money), as they will no longer have to go through (almost) every patch and evaluate whether it could be fixing a security problem. This means that such scenario will get a lot cheaper, while the defenders won't gain much from that, as one still needs to wait for the fix to be finalized and tested before deploying it in a production environment.
> My point was that if security patches are flagged as such from the start, it saves attackers lot of time (and money), as they will no longer have to go through (almost) every patch and evaluate whether it could be fixing a security problem.
Not really.
1. They can just check to see who made the commit - if it's a security researcher, it's obviously a vuln patch
2. The commits are obfuscated in hilariously obvious ways if you know what to look for
3. It's not that hard to look at a commit, it's kinda what they're paid for
> while the defenders won't gain much from that,
When the vuln is found a race begins between attacker and defender. The difference is that attackers know they're in a race and defenders find out two weeks later.
VFS/Page Cache/FS layers represent incredible complexity and cross dependencies - but the good news is code is very mature by now and should not see changes like this too often.
I'd like to add, for the less tenured developers around: "with the level of experience you're just going to live the fact that there'll always be bugs."
Experience have thought me, deal with problems when it is a problem. Dealing with could be problems can be a deep, very deep rabbit hole.
The commit message gave me the feeling that we should have just trust the author.
https://github.com/torvalds/linux/commit/f6dd975583bd8ce0884...
> I always try to shy away from making things look nicer
That's understandable, though from my experience, lots of old bugs can be found while refactoring code, even at the (small) risk of introducing new bugs.
Also, try to avoid small commits / changes; churn in code should be avoided, especially in kernel code. IIRC the Linux project and a lot of open source projects do not accept 'refactoring' pull requests, among other things for this exact reason.
Not sure what you mean by that
FWIW, I tend to err on the side of "do it", and I usually do it. But I have been in a situation where a customer asked for the risk level, I answered to the best of my knowledge (quite low but it's hard to be 100% sure), and they declined the change. The consequences of a bug would have been pretty horrible, too. Hundreds of thousands of (things) shipped with buggy software that is somewhat cumbersome to update.
> Experience have thought me, deal with problems when it is a problem.
Experience has taught me that disparate ways of doing the same thing tend to have bugs in one or more of the implementations. Then trying to figure out if a specific bug exists other places requires digging into those other places.
Make it work. Make it good. Make it faster (as necessary) is the way my long-lived code tends to evolve.
Anyone who doesn't, hasn't been burnt enough so far, but will be burnt in the future.
Not refactoring code also sacrifices long term issues in return for short term risk reduction. Look at all of the government systems stuck on COBOL. I guarantee there was someone in the 90s offering to rewrite it in Java, and someone else saying "no it's too risky!". Then your ancient system crashes in 2022 and nobody knows how it works let alone how to fix it.
We have a team like this -- their processes often failing, and their error reporting is lacking in important details. But they are not willing to improve reporting / make errors nicer (=with relevant details), instead they have to manually dig into the logs to see what happens. They waste a lot of time because they "shy away from making things look nicer."
— C.A.R. Hoare, The 1980 ACM Turing Award LectureIf the code is causing more work to maintenance and new development, sure it may make sense to refactor it. Otherwise, like the human appendix, just leave it alone until it causes a problem.
(from the downvotes it seems like some people don't want software to be a science)
And they work both ways -- there are anecdotes that making code beautiful leads to better outcomes, and there are anecdotes that having ugly code leads to better outcomes.
This means you cannot use lack of scientific research to give weight to your personal opinions. After all, that argument works in either direction ("There is no evidence that leaving duplicate code in the tree leaves to worse outcomes... A true science would provide...")
One of the prevailing features of well driven open-sources project is - you're encouraged to improve the code i.e. make it better {readable, maintainable, faster, hard}. You're not encouraged to change it for the sake of change i.e impress people.
I've the feeling it is the first case because it reduced the number of lines and kept source readable. Aside from that, I don't think good developers want to impress others.
By this logic why not write the entire kernel in assembly? Tools evolve and improve over time and it makes sense to migrate to better tools over time. We shouldn't have to live in the past because you refuse to update your compiler.
Otherwise you need to download a prebuilt compiler anyway, and whether that is C11 or C++11 is rather unimportant.
The GCC version requirement is 5.1 (which is 7 years old). Before that, it was 4.9, 4.8, 4.6 and 3.2. It has never been 4.7.
Use of newer versions of C than C89 which provides solutions to actual issues is perfectly fine. C11 was picked because it does not require an increase in minimum GCC version to use it, making your entire argument pointless.
The Linux kernel is already pretty lenient, as many alternatives have their a compiler on the tree and target only that.
Writing [repeated, if needed] monotonically increasing counters like this is a really good testing technique.
Luckily, almost all (if not just all) these millions of devices which will never be updated never ever received the vulnerable version in the first place. The bug was only introduced in 5.8 and due to how hardware vendors work phones are still stuck in 4.19 ages (or better, 5.4. but no 5.10 besides Pixel 6)
Based on the description, sounds like it should be quite possible.
I get responsible disclosure is important, but should we not give people some more opportunity to patch, which will always take some time?
Just curious.
Also, nice work and interesting find!
The announcement only serves to let the rest of the public know about this and incentivize them to upgrade.
It puts me, as a defender, at an insane disadvantage. Attackers have the time, incentives, and skills to look at commits for vulns. I don't. I don't get paid for every commit I look at, I don't get value out of it.
This backwards process pushed by Greg KH and others upstream needs to die ASAP.
(Thanks Max for handling this well and politely and for putting up with everyone’s conflicting opinions.)
I can't help but physically shake my head as I write this. I can't imagine actually asking people to try to play pretend security through obscurity because folks still can be arsed to implement some sort of reasonable update strategy. I have enough experience in tiny and huge shops to say that it's a matter of prioritization and it's just a blatant form of technical debt and poor foresight.
Via HTTP, all access logs of a month can be downloaded as a single .gz file. Using a trick (which involves Z_SYNC_FLUSH), we can just concatenate all gzipped daily log files without having to decompress and recompress them, which means this HTTP request consumes nearly no CPU. Memory bandwidth is saved by employing the splice() system call to feed data directly from the hard disk into the HTTP connection, without passing the kernel/userspace boundary (“zero-copy”).
Windows users can’t handle .gz files, but everybody can extract ZIP files. A ZIP file is just a container for .gz files, so we could use the same method to generate ZIP files on-the-fly; all we needed to do was send a ZIP header first, then concatenate all .gz file contents as usual, followed by the central directory (another kind of header).
Just want to say, these people are running a pretty impressive operation. Very thoroughly engineered system they have there.
Please update your devices ASAP!
Sorry!
Does it need patching? Of course. It’s not a privilege escalation remote code execution issue though, and even if it was, it would be on a tiny fraction of running devices right now.
That's correct and I misjudged the situation. Sorry!
*stares at kernel-3.10.0-1160.59.1.el7*
https://lists.debian.org/debian-security-announce/2022/msg00...
The security-tracker site is now updated as well.
21.10 appears to be lacking the patch.
It's always slower than we'd like.
Exciting for the implications of opening many locked down consumer devices out there.
Nightmare for the greater cyber sec world...
That doesn't sound right.
> if I recall correctly, you just chop off the GZIP header.
...to get the raw DEFLATE stream, that is. You still need to attach any necessary metadata for PKZIP, which Max mentions. Their approach for converting between the two is pretty clever: it's so elegant and simple that it seems obvious, but I never would have thought of it. Very nifty, @max_k!
What a supérioritégrandeur!
I mean, compiling 17 kernels alone takes so long that most people would've given up in between.
But yeah, incredibly impressive persistence.
Nah, that's the most fun part. Once you have one kernel that works and one that doesn't, you can be pretty sure that you'll eventually find the cause of the bug. The part where I would have given up is the "trying to reproduce" part.
The real deal was tracking it down and creating a reproducible test.
What are the memory savings of this splicing approach as compared to streaming [through userspace]?
If the application were processing the data in some way, then it would be worth it. Otherwise it is better to skip all of that work.
It was a random, intermittent file corruption that didn't cause real harm to the authors organization and was, clearly, very tricky to track down.
I’ve worked with (and been) a dev for several decades, and I can count on one hand the number of folks who would have a chance of figuring this out, and 2 fingers the number of folks who WOULD.
Of course, most never try to optimize or go so deep like this that they would ever need to, so there is that!
In both cases the researchers chose sort of punny names that were also self descriptive and obvious once you read how to produce the exploit. 'Dirty Pipe' is literally the recipe for this exploit / corruption. Maybe your name seems funny to you for some reason that isn't obvious / shared.
Really impressive debugging too.
The fix:
+++ b/lib/iov_iter.c
@@ -414,6 +414,7 @@ static size_t copy_page_to_iter_pipe(struct page \*page, size_t offset, size_t by
return 0;
buf->ops = &page_cache_pipe_buf_ops;
+ buf->flags = 0;
get_page(page);
buf->page = page;
buf->offset = offset;The bug is that they are reusing (or, repurposing) an already-allocated-and-used buffer and forgot to reset flags. This is a logic bug, not a memory safety bug.
In fact, this might be a prime example of "using Rust does not magically eliminate your temporal bugs because sometimes they are not about memory safety but logical". Before that my favorite such bug is a Use-After-Free in Redox OS's filesystem codes.
Pro tip for random HN Rust evangelist: read the fucking code before posting your "sHoUlD HAVe uSED A BeTTER lANGUAGE" shit.
You could argue that some languages distinguish raw memory from actual objects and even when reusing memory you would still go through an initialization phase (for example placement new in C++) that would take care of putting the object into a known default state.
See https://godbolt.org/z/Wh5KcTaGY for what I'm talking about, the local allocation is easily eliminated by the compiler.
The equivalent in C is to create a temporary local variable with an initializer list then write that variable to the pointer.
edit: the "struct pipe_buffer" in question[1] has one field that even the updated code doesn't write - "private". Not sure what it's about, but it's there. Not writing unrelated fields like that is probably not much of an issue now, but it certainly can add up on low-power systems. You might also have situations where you'd want to write parts of an object in different parts of the code, for which you'd have to resort to writing fields manually.
[1]: https://github.com/torvalds/linux/blob/719fce7539cd3e186598e...
Re: your other points, "reusing a pre-allocated struct from a buffer" is basically object initialization, which is different from other times you want to write fields. In general an object initialization construct should be used in those cases, this whole thread being an argument why. Out-of-band fields such as the "private" field are a pain I agree, but they can be separated from an inner struct (in which case the inner struct is the only field that gets assigned for initialization).
Taking a step back, the true solution is probably to have a proper constructor... And that can be done in any language, so I'll stand corrected.
To be clear, some base level of trivially safe code is certainly a nice-to-have. I just don't think the amount that helps for a kernel is that much, and the added boilerplate on more unsafe things might even obscure issues.
And there are recent plans to move the kernel to a newer C version anyway.
This statement is incorrect. They are using an arena allocator, and there is no way for it to know if it is reusing one of the elements or using that element for the first time. To do this in Rust you would probably be using the MaybeUninit type: https://doc.rust-lang.org/std/mem/union.MaybeUninit.html
However, you are partly correct. In Rust, when using the MaybeUninit type, it is still possible to partially initialize an object and then return it as if it were fully initialized without hitting a compile error. https://doc.rust-lang.org/std/mem/union.MaybeUninit.html#ini...
If you do the whole struct at once, rather than one field at a time, then the compiler still has your back:
let foo = unsafe {
let mut uninit: MaybeUninit<Foo> = MaybeUninit::uninit();
uninit.write(Foo {
name: "Bob".to_string(),
list: vec![0, 1, 2]
})
}buf = &pipe->bufs[i_head & p_mask];
i think it is a problem with merging and needing to reset the flags, but i didn't want to waste too much time really trying to find the root issue and wtf merging is/why do it.
I'm not a fan of such a policy. That usually leads to people zero-initializing everything. For this bug, this would have been correct, but sometimes, there is no good "initial" value, and zero is just another random value like all the 2^32-1 others.
Worse, if you zero-initialize everything, valgrind will be unable to find accesses to uninitialized variables, which hides the bug and makes it harder to find. If I have no good initial value for something, I'd rather leave it uninitialized.
So use a language that has an option type, we've only had them for what, 50 years now.
The issue is with dynamic memory allocation as that would be the responsibility of the allocator (and of course the kernel uses custom allocators).
What bothers me about the Linux code base is that there is so much code duplication; the pipe doesn't use a generic circular buffer implementation, but instead rolls its own. If you had the one true implementation, you'd add those annotations there, once, and all users would have it, and would benefit from KMSAN's deep insight.
Every time I hack Linux kernel code, I'm reminded how ugly plain C is, how it forces me to repeat myself (unless you enter macro hell, but Linux is already there). I wish the Linux kernel would agree on a subset of C++, which would allow making it much more robust and simpler.
They recently agreed to allow Rust code in certain tail ends of the code base; that's a good thing, but much more would be gained from allowing that subset of C++ everywhere. (Do both. I'm not arguing against Rust.)
But I think it would be troublesome to use such a hypothetical feature in C if it's only available in some compiler-specific dialect(s), because you need to coerce to any type, so it would be hard to hide to hide behind a macro. What should it expand to on compilers without support? It would probably need lots of variants specific to scalar types, pointer types, etc., or lots of #if blocks, which would be unfortunate.
Zig is a nice language with this feature, and it fits into many of the same use cases as C: https://ziglang.org/documentation/0.9.1/#undefined
I'm a C++ guy, and the lack of constructors is one of many things that bothers me with C.
This proposal sounds great until you find out that this is a hard problem to solve reasonably well in the compiler and no matter what you do there will be valid programs that your compiler will reject.
> valid programs that your compiler will reject.
Do you mean "valid programs where everything is initialized when an object is created will somehow fail to detect that"?
Valid programs where a field is left uninitialized at creation time, but the programmer makes sure it's initialized before it's used.
I agree that it’s a hard problem, but I don’t think that language designers should use that as an excuse. I think that language designers should really double–down and require that every field be initialized, otherwise people will just forget. The code will be carefully written at first, but then someone new will come along and add a new field, and then the existing code is insufficient. With no errors from the compiler there is no way to ensure that the new programmer updates everything to accommodate the new field.
Rust does pretty well in this regard, but you can still use unsafe blocks to get a partly–initialized struct if you really want one. I like Rust, but I wish that it went all the way and didn’t allow it at all.