So You Want to Rust the Linux Kernel?
paulmck.livejournal.com
paulmck.livejournal.com
There are plenty of other kernels available for embedded work, and my understanding of Rust in the Linux kernel is that it's mostly drivers, which seems like it would be optional for your use case.
That said, I've not done much embedded, I may be misunderstanding something.
The parent comment is complaining that Rust will make kernel dev too hard? Are they joking? It is already complicated.
I would rather see a more formally constructed language used for the kernel than risk adding more C-related bugs.
And specifically, one of the ways in which it is complicated is that it's mostly done in a language that doesn't natively enforce memory-safe constructs, forcing developers to walk a best-practices and tooling tightrope to avoid writing code that breaks.
And to call Rust's build system is laughable. Issuing `cargo build` could not be simpler (and it includes automatic downloading of dependencies, too).
For example, how many bytes are in an `int`? What is the endianess of your 32bit `int`?
Rust puts up front a lot of the information crucial for cross platform development of hardware drivers.
Rust has few undefined edges which is exactly where C gets into trouble. C, by design, has a bunch of undefined edges.
Honestly, the biggest downside to Rust in the kernel is that Rust is backed by the LLVM, which doesn't have as many supported targets as GCC does. (And, AFAIK, GCC is generally the preferred compiler for the kernel).
Fortunately, there are now two competing approach to make GCC a viable compiler to Rust[1], and it's moving really fast. I think the LLVM-only situation won't get in the way more than a few more years.
[1]: https://lwn.net/SubscriberLink/871283/c437c1364397e70e/
Rust and Linux both have the same name for say the 32-bit unsigned integer type, they both call that u32. This is of course a coincidence, and not an unlikely one, but it probably doesn't hurt for Linux people getting familiar with Rust.
also, being able to use linux in an embedded system as a resource constrained developer is hugely important today. adding more tools means more dev cost.
Ouch. Good luck with that.
But "I don't want to learn something new"? No, that's not a legitimate reason.
But Christ almighty, you don't get to block progress of all humanity (and Linux is a humanity scale project now) because you're too lazy to learn some new syntax. Memory safety is important. Memory safety bugs cause billions of dollars of damage every single year. The project if solving them is too important to get tangled up in the weeds of people who don't have "resources" to RTFM.
fn sum_odd_numbers(up_to: u32) -> u32 {
why fn? why switch the order around? fn and u32 and all that as if keystrokes are such a scarce resource, yet u32 f(u32 a) saves so much more than fn f(a: u32) -> u32.Closures use pipes instead of any of the existing syntaxes we're used to.
Why couldn't they just use the familiar concept of classes instead of whatever is going on with their strange NIH traits? NIH describes the whole language, every little thing has to be different somehow.
C with Rust's safety guarantees, OOP style classes, and a few more bells and whistles could have taken over the world a lot faster, but instead we have something with a high friction to learning that will be adopted much more slowly by either programmers with a fresh start to whom every language is equally weird, or the relative minority of experienced programmers with the will and free time to push past the friction.
The people pointing out problems with Rust have more experience as developers in their little fingers than Rust's internet fan club has in their whole bodies combined.
In a mature, intelligent discussion, if you have a point, you make it. On HN, if you have a point, you say "there are reasons; I'm not going to give them; also, you're a newbie".
It's comments like yours that make HN unbearable nowadays.
It's not that "modern" languages are doing something new. Everybody was doing it like that, besides the "C language family". They're the outlier, not the other way around.
Parsing! it is faster to parse and probably uses less memory too. This is not only important for compile time but also for editors with code suggestions.
C has a relative long compile time. If you look at golang, which has a goal of a fast compile time, you will see similar syntax.
PS.
Note that Ken is one of the fathers of both C and golang.
Parsing is ridiculously fast nowadays. Compilers spend far more time in phases after AST construction (like register allocation, SSA transforms, various vectorizaton passes) than they do constructing their ASTs. Parsing speed is a nonsense argument these days. It's the same level of rigor as painting a car red to go faster.
> If you look at golang, which has a goal of a fast compile time, you will see similar syntax.
Rust critics: "Some of Rust's language choices are clearly inspired by Go and may not have been good ideas"
Rust people: "We didn't copy Go! We independently through a rigorous analysis arrived at our syntax!"
Also Rust people: "If you look at golang, which has a goal of a fast compile time, you will see similar syntax."
> Note that Ken is one of the fathers of both C and golang.
Creating one of these things makes Ken made imperfect. Creating both of these things makes Ken a menace.
Some of these things have tradeoffs, but I think once you get to grips with what the borrow checker is asking of you (doesn't take long), it's a pretty easy language - certainly simpler than C[0], far simpler than C++.
[0]: Non-buggy C, that is. I wrote a pretty basic program in C the other day (I'm not a C programmer) and I literally spent 80% of the time looking at output from -fsanitize=address to catch stupid off-by-one errors.
Who is “we”? Pipes are one of the existing closure syntaxes I was used to pre-Rust. And, I mean, I’m probably not alone: lots of current and ex-Rubyists around.
It has benefits. It's easier to parse for both humans and machines and it allows for easier type inference.
>Closures use pipes instead of any of the existing syntaxes we're used to.
Closures in Ruby use pipes. It's a common syntax.
>Why couldn't they just use the familiar concept of classes instead of whatever is going on with their strange NIH traits? NIH describes the whole language, every little thing has to be different somehow.
Needless to say, traits aren't NIH either, and there's good reasons for avoiding class-style polymorphism in a language like Rust.
Here's some reading to help explain. https://stevedonovan.github.io/rust-gentle-intro/object-orie...
It also happens to include this choice quote:
>I once attended a Java user group meeting where James Gosling (Java's inventor) was the featured speaker. During the memorable Q&A session, someone asked him: "If you could do Java over again, what would you change?" "I'd leave out classes," he replied. After the laughter died down, he explained that the real problem wasn't classes per se, but rather implementation inheritance (the extends relationship). Interface inheritance (the implements relationship) is preferable. You should avoid implementation inheritance whenever possible
And that's basically what Rust gives you
For example, suppose you just invented the type Qux and you always know how turn any Baz into a Qux.
Rust's standard library provides four interesting conversion traits that may be applicable for different purposes, From, Into, TryFrom, and TryInto. Oh no, implementing all of them sounds like a lot of work and who knows which anybody would need?
Just implement From<Baz> for Qux
A blanket implemention of Into<U> for T when From<T> for U gets you Into, then a blanket implementation of TryFrom<T> for U when Into<U> for T gives you TryFrom, and finally a blanket implementation of TryInto<U> for T when TryFrom<T> for U gets you TryInto as well.
So you only wrote one implementation but anybody who needs any of these four conversions gets the one they needed.
Another more obvious example is the blanket implementation of IntoIterator for any Iterator. This makes it easy when you want to be passed a bunch of Stuff you're going to iterate over, just give a trait bound on your parameter saying it has to be IntoIterator for Stuff. You needn't care if the user of your function passes you an array, a container type, or an iterator, since all of them are IntoIterator.
It would have been so easy to overlook this, and end up with a language where people find themselves collecting up iterators into containers over and over to make more iterators, or else constantly turning containers into iterators to call functions.
It is true that SFINAE means you can write very general templates in C++ and just not worry about whether they compile with any particular parameters, which is superficially similar to a blanket implementation of a Rust trait. But, because Rust's traits have semantics not just syntax the actual effect is different.
And that makes the blanket implementations very comfortable in Rust, whereas the equivalent pile of enable_if templates in C++ would be trouble.
Remember C++ comes with template language that is explicitly labelled as having unenforced semantic value. It's for documentation, a human reading them can see for example that you expect this parameter type to have full equivalence. However the compiler only cares about syntax, and the syntax just says the type has an equals operator.
This is a great contrast with Rust, whose traits have semantics and so core::cmp::Eq and core::cmp::PartialEq are different traits and the compiler cares which one you asked for even though the syntax is the same.
Don't concepts give C++ the same semantic power? What am I missing? Concepts give C++ generic code exactly the same nominative constraints you're describing, yes?
Besides: even without concepts, you can use tag types and traits to similar effect. Not everything needs to be duck-typed, even in C++17!
One of my main problems with using Rust instead of C++ is losing metaprogramming power. I really like the template system.
Not really. But, it's on purpose, and, in their context it makes sense. These aren't idiots, they know there are however many bajillion lines of C++ code and if Concepts doesn't work with that code it's useless.
In Rust, you have to explicitly write implementations of traits. If the author of stupid::Bicycle doesn't implement nasty::Smells and the author of nasty::Smells didn't implement it for stupid::Bicycle either, then a stupid::Bicycle just isn't nasty::Smells.
Even "marker traits" where the implementation is empty, are still explicit. If you have a type that implements Clone but you didn't say it is Copy, then it isn't Copy, even though maybe the type would qualify as Copy if you had asked. Rust will always move this type, because you didn't say copying it was OK.
This is where the semantic strength comes from, and it was fine in Rust because Rust "always" did this. There isn't a bajillion lines of Rust code that doesn't know about traits, and so it's OK to require this.
In C++ they realistically couldn't do that. So, next best thing is to say that a Concept is modelled on syntax. You mentioned duck-typing, and that's what concepts are assumed to be for.
Now, of course you can use tag types, for a home grown Concept it actually might be viable, if you've only got two users and they both know you personally you can just tell them to use the tag type and they'll both add it. But for the C++ standard library Concepts that was not practical. So, that's not how Concepts is being taught or has been used in practice, unlike Rust's traits.
It's also a shame that your comment is being downvoted into oblivion. I remember when HN was for discussion, not for karma-enforced orthodoxy.
The C++ Standard committee disagree with you - and not for lack of trying, either. The C++ Core Guidelines are the best they could come up with, and these are not rigorously enforced unlike the safe subset of Rust.
I hate that paper. People misrepresent it all the time.
[1] https://docs.google.com/document/d/e/2PACX-1vSt2VB1zQAJ6JDMa...
A lot of the "whacky, quirky" stuff is to avoid the clusterfuck that is C and C++ parsing.
People want IDEs. They also want their IDE to do smart things. Having to compile the universe to get the IDE to do something smart doesn't work very well.
Take a "simple" task like "Is this is a variable name? That is a type definition?". C and C++ have to practically do a full compile to figure that out (see: typedef and templates). Rust, on the other hand, just has to read a few tokens and it knows definitively the answer to that question.
And, you may call this stuff "quirky", but practically all the languages designed in the last 10 years have converged to similar goals and syntaxes to Rust. Language designers know that people don't like "different from C" but they do it anyway.
Think about whether your arguments are stronger than those who are designing new programming languages that they want adopted.
There's a crucial part missing: the investment [...] to build (memory) unsafe drivers.
I find it very troubling when memory unsafe language programmers don't weigh memory safety at all in their views.
(note that this doesn't necessarily refer to Rust; there's people who swear that one can build drivers also in Golang, so assuming that is realistic, the same reasoning applies)
C-style syntax and community interest also favor Rust.
This spec argument has always seemed like a red-herring to me. Can you explain why the Ada spec would be a significant factor in this instance?
I think GC is optional with Ada, as far as I know the memory safety comes from raising exceptions (or refusing to compile) when it detects memory-unsafe operations (array bounds checking etc).
>C-style syntax and community interest also favor Rust.
That's fair, C-like languages are instantly familiar with software people and rust seems to have a more hip image than old man fuddy-duddy Ada.
>This spec argument has always seemed like a red-herring to me. Can you explain why the Ada spec would be a significant factor in this instance?
I'm not arguing for Ada over Rust (I'm not a software guy) so I don't mean it as a red herring, but wouldn't a suite of static verification tools and a formally verified compiler require a spec to be tested against?
I think a spec is often a red herring because a bunch of folks living in the slum of C, when asked if they would all like to move into Rust's nice 3 bedroom by the park, instead always seem to ask: Wouldn't it be better to form a committee about building us a cathedral? Ada might be better. Someone should try it, but until then I'll take Rust.
Remain interested in the potential advantages of Ada compared to Rust for Linux kernel development, if you would care to point me in the right direction.
Ada is safe compared to the other languages of its day, but I'm not sure it compares favorably with Rust. IIRC it does have more of a focus on safe arithmetic rather than memory safety.
Ada shines in its specification power, how developers can express what the code is supposed to do (strong typing, ranges, contracts, invariants, generics, etc.). And then you can either check your code at "compile time" with SPARK [1], that provides a mathematical proof that you code follows the specification. SPARK also proves that you don't have buffer overflows or division by zero for instance. Or you can have checks inserted in the run-time code which greatly improves the benefits of testing as every deviation from specifications will be detected, not only the ones you decided to check in your tests.
In terms of memory safety, Ada always had an edge on C/C++ because of the lower usage of pointers (see parameter modes [2]) and the emphasis on stack allocation. Now with the introduction of ownership in SPARK it's getting on par with Rust on that topic.
[1] https://learn.adacore.com/courses/intro-to-spark/chapters/05... [2] https://learn.adacore.com/courses/intro-to-ada/chapters/subp...
there was also a lot of unchecked conversion under the hood.
(disclaimer: I haven't really paid attention to newer versions of Ada)
(Safe) Rust is Data Race Free. You can mistakenly write a data race in Go, just as you could in C, and, even though under normal circumstances Go feels like it's memory safe, once you've got a data race all bets are off and you potentially no longer have memory safety. In (safe) Rust you can't write a data race, so, this can't happen.
If you write simple serial programs, you don't care, your programs never have data races because they have no concurrency and therefore can't possibly have concurrent memory writes. But obviously Linux is not a serial program.
We proven time and time again we can't write safe c/c++ code from a security perspective. I don't think we've proven that crashing now and then to fix our data races is a bad outcome. There are always bugs to be fixed.
The big problem is SC/DRF. Most languages promise (at best) your program is sequentially consistent if data race free. And it turns out we can't debug non-trivial programs unless they're sequentially consistent because it hurts our brains too much.
Causality is something we're all used to in the real world. The man hits the ball with the bat, then the ball's trajectory changes, because it struck the bat. Sequentially consistent programs are the same. X was five, then we doubled it, now it's ten.
But without sequential consistency that's not how your program works at all. The ball goes flying off over the wall around the field, then the pitcher throws the ball, the ball lands in the catcher's glove, the umpire calls it, the batter hits the ball, and yet at the same time misses the ball... We doubled X, now it's five, but before we doubled it that was nine or sixteen? What?
It must be noted, though, that data races can lead to memory safety issues in languages which are otherwise free of them, Go being a prime example of that: the built-in hashmap is not thread-safe, and does exhibit memory corruption issues under unprotected concurrent writes and accesses (only concurrent unprotected reads are OK).
There is a “best effort detection” of concurrent misuse, but it’s just that.
Second, this is exactly the overconfidence that I find troubling. Vulnerabilities caused by memory unsafety are real, pervasive, and inevitable.
A few resources you should look at:
- the official reports/post by Mozilla and Microsoft about the percentage of vulnerabilities caused by memory unsafety
- the mailing lists with the Linux kernel issues (I don't imply any severity; just the amount)
There are many summarized posts for the lazier, and I'm talking about high profile authors (or authors writing on behalf of high profile companies). Here's a semi-random one:
https://msrc-blog.microsoft.com/2019/07/18/we-need-a-safer-s...
I'd assume naively that such a construct would, of necessity, involve the needs of hardware devs.
I was emphatically not suggesting that the needs of hardware devs should be ignored, just that those needs should not constrict kernel development to the detriment of other concerns.
* - There is some cryptography stuff in the kernel (that probably doesn't need to be in there anyway) and that isn't written by hardware devs I suppose.
The actual cognitive cost does not come from Rust itself, but from mixing the two memory models and interop between them, IME and is probably a bad idea in general.
Should just boil the ocean and write the entire LK with Rust only.
> Including some wild speculation about how Rust's ownership model might be generalized
I recommend starting with the latests post in the series and the conclusions(https://paulmck.livejournal.com/64209.html and https://paulmck.livejournal.com/65056.html).
From the snippets at the beginning, it looks like the author documented how they started learning about how to solve these problems in Rust, and it does not take too long until the comment section starts pointing the author at crossbeam-rs, GhostCell, etc. which only start appearing in the series in the last couple of sections.
Learning Rust by implementing a linked-list is pretty hard, but "Learning Rust by implementing RCU" has to be, by far, the hardest way to learn Rust that I know of. Kudos to the author for pushing through this so quickly and thoroguly. The posts are all excellent.
I wish the author would go a little bit more into crossbeam-rs. One thing the author does not consider, is that a safe Rust interface to RCU is essentially a "proof of correctness" for RCU, i.e., it proves that using RCU via that interface cannot introduce undefined behavior. (Or maybe more technically, the contracts on the correctness of unsafe code within the safe wrapper provide a potentially incorrect "proof" of this).
AFAIK no such correctness proof for the C APIs exists (there are a couple of linters that catch some issues, but there is no proof that say that if you stick by those warnings your code is UB free).
The author seems to be assuming that the C APIs to RCU have been proved correct, which AFAIK is incorrect, and by ignoring the safe Rust APIs to RCU, the author is missing out on potentially finding issues on the correctness of the C APIs that Rust would discover.
Basically, the question: "Can we expose Linux RCU apis to Rust?" is definetly worth asking, but the question "Can Rust prove Linux RCU APIs correct?" is much more interesting. If it can't, then why not? The fact that we have such proofs for other RCU and hazard pointers APIs in Rust shows that this is at least possible in general. And if Linux RCU can't be exposed in safe Rust, precisely understanding why would definitely be illuminating (Can Linux RCU be proven correct outside Rust, but not in Rust? Or maybe does Linux RCU APIs can actually be proven incorrect by Rust?).
You can implement doubly-linked lists in safe Rust by using Rc and Weak, Option, and RefCell (or Arc and Mutex for a safe concurrent doubly-linked list).
Learning about safe Rust is gentler and more useful for a beginner, than starting to learn Rust by learning about "unsafe code".
You have to know what safe Rust allows so that you can appropriately know to which contract unsafe Rust must adhere to.
You can, but not without introducing runtime overhead relative to C and C++, e.g., by forcing reference-counting where none would be necessary in C.
Rust's safety checks do in fact block some safe and zero-overhead abstractions familiar to people working in C and C++, and denying that isn't helpful. Stating that these techniques can be implemented in Rust with overhead is missing the point.
Arguing that the added overhead is "minimal" (which is a meaningless word, since the word "minimal" just means "I don't care about your scenario") is still missing the point: it's overhead that's not present in unsafe languages.
The reason people use languages like C++ and Rust is to get zero cost abstractions and explicit control over the machine. If you want good performance but don't care about precise control over the machine, write Java or C# or something high level like that. You'll be just as safe as you would be in Rust and more productive.
I'm really tired of this motte and bailey stuff. Rust proponents say Rust gives you fine grained machine control and safety with no performance compromises, and then when you point out that the language doesn't quite live up to that promise, Rust proponents start telling you that you didn't really need that performance anyway. This argumentative tic is annoying.
The runtime overhead is minimal. 96% of it comes from the pointer indirection in the linked list, and this you have either way.
Last time I benchmarked, the checks made the list 4% slower, which is acceptable for many, since doubly-linked lists are already very slow.
This 4% is buying you a correctness proof that your list API cannot introduce undefined behavior. If you write this doubly-linked list abstraction in safe Rust, and make a mistake, the resulting code will still not have undefined behavior.
C and C++, even paying this 4% performance cost, do not provide this guarantee, which is valuable to many.
Unsafe code that removes this 4% runtime overhead is an optimization.
The fact that you can remove this overhead by providing a safe wrapper over unsafe code is a selling point of Rust, but learning how to write such optimized code without learning first what it takes to make it correct IMO completely misses the point of Rust.
It is very easy to write broken safe Rust abstractions over broken unsafe code. At that point, you are in a worse place than C or C++, since you are paying many costs for introducing Rust in a project, but leaving the most valuable feature off the table (correctness, lack of segfaults, lack of data-races, etc.).
Also linked lists are often [1] used in C and C++ for intrusive chaining of nodes otherwise owned by other data structures and the reference counting would either impose overhead to those other datastrucutres or be pointless.
[1] In fact I would say this is very common at least in C++ as linked lists make otherwise very poor datastructures by themselves.
A doubly-linked list is already space inefficient, it adds 2 pointers per element.
Using Rc + Weak adds one extra word per element to store the two (strong and weak) refcounts. So for a 64-bit type, a Vec takes 1x space 1xN, a list takes 3xN, and a ref-counted list 4xN.
Performance wise, the extra space doesn't change anything, since that word is allocated with the element in the same cache line and would have been read anyways (even if you don't access it, the hardware does access it).
If you are using a doubly-linked list, you really don't care that much about any of this, and the only thing that makes a real difference, is the guarantee that your list implementation is correct.
You can still use double linked lists in high performance code, as long as the list traversal doesn't happen on the critical path (or at all).
"Performance is almost the same in real world cases"
"You don't really need that data structure anyway"
"You can learn how to optimize code later"
People come and want to use Rust for some fast code. They know what they want the code to do and want it to be as fast as the code they wrote before, and see how Rust would help them achieve that. The above statements are then wrong or patronizing or completely misses the point.
That doesn't matter if a solution with good cache locality doesn't exist for your problem.
Why would you want to use a non-intrusive doubly-linked list?
AFAIK there is only one answer: you have very big objects and you need O(1) splice.
If you don't have very big objects, or you don't need O(1), pretty much any other data-structure in existence is going to give you much much better performance than a doubly-linked list.
For the only use case for which non-intrusive doubly-linked lists are good at, however, there exists no hardware you can buy for which you can measure a performance difference between ref-counting and not ref-counting.
This has nothing to do with Rust. The same applies to C++, or Java, or Python, or even C. You can do ref counting on any language, and all these languages run pretty much everywhere.
The only thing that Rust gives you over Java, Python, or C here is the possibility to implement a doubly-linked list in safe Rust that will perform the same but that the compiler will prove for you to be memory and thread safe.
Sure, Rust also allows you to use unsafe code, and then prove that safe and create a safe wrapper. It even does this for you and provides this in std::List. But what value does this add? It's more work, and it doesn't perform better, and why a significant number of people have argued in the past that adding std::List was a mistake in one form or another.
Or you have a list of objects with many references to parts inside the list and you want to rearrange those parts from those references at O(1) speed.
Lets say I have pointers two elements A and B in the list. I now want to say that B should come after A, I can do this easily in O(1) time with this with no overhead and all references sees this update instantly. Lets say I put integers in these so the data is embedded in the list. You have many other places referring to elements inside the list and you want those references to track the objects position.
How would you solve this using a Rust compatible data structure?
> The question is fair.
No it is not. Experienced people almost surely knows their use case better than you do.
> there exists no hardware you can buy for which you can measure a performance difference between ref-counting and not ref-counting.
Are you kidding me? Seriously, I don't see how you could believe this unless you never tried to use ref counting to replace other code and compare the performance difference.
You can do this with a vector of pointers as well, at lower space complexity cost (1 pointer per element instead of two), and a lower runtime cost (1 swap of two pointers instead of swapping 4 pointers).
Such a data-structure has also other advantages, like higher cache efficiency, etc.
With a vector of pointers to the elements (C++ std::vector<T*>) this can be easily done in O(1) as I explained above.
If now that I've proven you wrong, you want to change the problem to something else, feel free to state your new problem, and I'll proceed to prove you wrong again.
Putting B after A is not swapping them, it is putting the element B so it is right after A. You could technically interpret it as you did, but only a contrarian would do that. If you want more people to support rust then you should stop being a contrarian. If you want me and others to assume that the Rust community is full of hard to work with people then please continue, but that wont make people more likely to pick up rust.
If instead of behaving like you do here people would just say "Sure thing, in order to get that performance in rust you just do X and Y!" I bet people would be way more supportive of rust. But if working in the language means that people like you will come and argue like this then why not just write a C library, and then someone will write a Rust wrapper and people are happy? While if they'd write it in unsafe Rust people would come and complain like hell.
Or you could have been more clearer.
I didn't intend any animosity, but you clearly do.
> If you want more people to support rust
I don't care whether more people support Rust or not. I am just fighting the spectacular amount of disinformation in this thread by malicious actors that have never used Rust.
> Putting B after A is not swapping them, it is putting the element B so it is right after A.
Then you should have written that instead. Take the pointer in the vector right after A, and swap it with B. That way, A is now right before B, and B is now right after A.
This is a joke, right? I don't see how it can't be.
But in case it isn't, in almost all problems like this it is assumed that you don't want to change the order of the other elements.
For example, if you wrote a function MoveToDirectlyAfter(a, b), and it did move elements other than a or b people would say that your function contains bugs or that it doesn't do what it says.
Let's say a user reorders documents with drag and drop. What's the low-level equivalent? Could you help me come up with an algorithm that depends on this?
If write speed is important, I'd just use an in-memory LevelDB-like thing (so a log structured merge tree).
If the data is not much, then I'd optimize for code simplicity.
https://en.cppreference.com/w/cpp/algorithm/rotate
Is that what you want?
I think more what's more going on is that longtime Rust programmers can get frustrated by the sheer volume of people who have never used Rust but dismiss it because they hypothesize that they will need to regularly write things like linked lists in Rust just because they are the easiest data structure to write in C.
When in reality let alone the amount of high performance data structures in the standard library and crate ecosystem mean it's extremely unlikely they'll even feel to write any data structures in the first place. Programming Rust and C or C++ is a very different experience in numerous ways and it is unproductive to expect that not to be the case.
And correct use of linked lists is not uncommon at all. Binary trees are linked lists for example, would you implement binary trees as a vector pointing to vectors rather than a linked list containing pointers to similar linked lists?
I think this perspective you're espousing is "sometimes you just want to knock out a simple bespoke binary tree implementation", but I just don't think that use case makes sense. Either you want to use one from a good library or you have a really good reason to roll your own and it is worth the effort to do it well.
There is no library for generic tree data structures, you have to implement those yourself. There are libraries implementing a table interface using ordered binary trees but that is just one of many use cases.
This doesn't invalidate your point about there being use cases for binary trees that may necessitate you to write your own, I take you at your word that you've often had strong use cases for this, but I do think it illustrates something I believe to be true, which is that people reach for linked data structures too often, when contiguous ones are a better default.
If you don't have a use case in mind, I'd suggest just trying to actually use Rust for something instead of spending time trying to construct problems where Rust would be bad at the solution.
It might be possible that you can write nice versions etc, if so I'd like to see them, but it seems everyone is hell bent on just saying "You shouldn't do it!".
A beginner learns this in the first week:
struct Node<'a, 'b> {
left: Option<&'a Node>,
right: Option<&'b Node>
}
Many people don't know how to read or write, but you don't hear them every day claiming that "therefore it must be impossible", yet here we are.> but I don't see why I'd learn a language where I'd be forced to make those work arounds rather than use the version that fits my problem best.
> but it seems everyone is hell bent on just saying "You shouldn't do it!".
The claim was and still is "there are infinitely more effective - as in faster - ways to learn Rust than by starting with how to write doubly-linked lists".
This whole thread is you failing at 3rd grader logic hard and concluding:
- therefore it is impossible to write doubly-linked lists in Rust
- therefore doubly-linked list in Rust are efficient
- therefore it is impossible to learn Rust
- ...and many other things...
I've taught Rust to a 9th year old, but I don't think I could have taught Rust to a 2 year old, so arguably, I think you are right, in that probably _FOR YOU_ it is impossible to learn Rust, you wouldn't be able to grok how to write doubly-linked lists in Rust properly, you wouldn't be able to master Rust to write efficient code in it, etc.
That's ok, don't be frustrated by it, everybody is different.
Just try to stop projecting your limitations to other people. Particularly in this forum, or when talking with Linux kernel developers, were most people are smart experienced programmers that would learn Rust well in a day and master it in a week.
It could, but since the cachelines are adjacents the prefetcher will pull subsequent cache lines anyways, and since the objects typically used in doubly-linked lists are very large (much larger than 64-bit), then this doesn't matter on any practical application that I am aware of. Our results of 4% performance increase is for the worst-case that we found. Usually is 1% or less.
If you have a real-world benchmark that shows otherwise, can you please share it?
AFAIK there does not exist any hardware for which the performance model for:
- non-intrusive doubly-linked lists,
- typical object sizes used on these lists (>> 64-bit)
- typical operations done on these lists (splice, etc.)
would predict a performance difference, and in our case, having replaced unsafe lists with safe list in many large applications, we never were able to measure a difference in application performance, only in the synthetic worst-case micro-benchmarks.
So I am truly interested to learn about your use case.
The Linux kernel is made by people who spend weeks to shave one bit off the size of a data structure. Do you really think they'd respond positively to your saying they don't need to do that?
You really need to stop assuming what kind of computers other people are programming
The larger point here is you don't get to decide whether other people's use cases are legitimate. Either Rust gives you complete control of the machine or it does not. It does not, not in safe mode, but Rust people keep doing this annoying motte and bailey thing about it.
If they do need to shave a single bit off, they can do that, that's a neat feature too.
No, I am not. This conversation is exclusively about non-intrusive doubly-linked lists.
If you want to have a conversation about intrusive doubly-linked lists we can have it, but it is a very different data-structure.
> You really need to stop assuming what kind of computers other people are programming
Every week I touch x86, arm, ppc64, riscv and gpu assembly.
I write code for apps daily that run from 16-bit micro controllers to the largest super computers in the world, going through phones, desktops, cloud, etc.
I am not assuming what others are programming, but rather talking about what I program for, which include most hardware that anyone can buy today, and quite a bit of hardware that almost no one can even buy, as well as hardware that's not for sale.
...unless you use the XOR trick to store two pointers in one field?
This culture is why rust will never replace C++ and why C++ will never replace C. You can write the same computations in C++ as in C, and in Rust as in C++, but it isn't "idiomatic" so people are really afraid to do it because it will get harshly rejected by the community.
Your claim that you need to learn all low level details first is also not true.
There are ~20 million programmers in the world, and about ~16 million of those are Javascript programmers. Many of them don't know and don't need to know the difference between the stack and the heap.
Many of them regularly optimize their Javascript hotspots by re-writing them in Rust and compiling it to webassembly. And almost all of them benefit from Javascript libraries that do this internally.
Rust is for many of them the ramp up into lower-level systems programming, and this is one of the reasons the Rust project has so many contributors. Rust enables Javascript programmers to actually hack on the Rust compiler, Firefox, etc.
This is something that C and C++ never achieved, and one of the main reasons for the Linux kernel to want to use Rust (they want to attract more junior developers to increase the developer base and make the project more accessible).
Your tone that the large majority of programmers in the world are somehow "doing programming wrong" by learning higher-level and safe languages first, and delving into low-level details as they need to, sounds very elitist and I personally find it disgusting.
My point was that Rust wont replace C and C++ as long as the Rust community is as it is now. And from your comments it seems like Rust isn't intended to replace C or C++, but act as a low level language for people who don't know how to write low level code, like javascript programmers.
> Your tone that the large majority of programmers in the world are somehow "doing programming wrong" by learning higher-level and safe languages first, and delving into low-level details as they need to, sounds very elitist and I personally find it disgusting.
I didn't say that everyone has to learn rust by learning how to write the low level parts. I am saying that if someone wants to learn how to write low level parts of rust rather than the safe parts because they are used to writing performance critical bits of code in C or C++ then you shouldn't discourage them from doing that. Please don't put words in my mouth.
It is fine for you to not be able to think of any scenarios where the battle tested standard libraries wouldn't be sufficient. But as someone who wrote low level libraries running in production at Google that doesn't apply to me, for me knowing what it would take to eek out performance in Rust compared to C++ is absolutely necessary and given how hostile people are my guess is that the Rust wouldn't look pretty so there is no reason for me to even try to pick up Rust.
And you can't say that I am not your intended user, if you chase away people like me from the rust ecosystem then there is no way that Rust can replace C++ in any reasonable timeframe.
What I'd like to see is a lot of examples and such showing how to write very complex things using unsafe rust. Because without that there is no way rust can compete, there are already plenty of people who know how to write highly performant code using unsafe C++.
I don't at all think that there are not "any scenarios where the battle tested standard libraries wouldn't be sufficient". I spent a number of words on exactly that in another comment below. My point in that comment was that the use case I don't believe actually exists in production code is "I'm just gonna knock out a simple bespoke implementation of a binary tree". I'm not saying people don't do that, but they shouldn't. If what you need is simple, there is a good library for it. And if the library for it is not sufficient, then it is important enough to invest time into writing a good implementation.
Maybe here's a way to put it: simple bespoke implementations of linked data structures are easy in both c++ and rust using ref counting, slightly better but low-investment implementations are still pretty easy in c++ but are not really possible in rust, and really good implementations are possible in both languages if the investment is worthwhile for the use case. It is true that you would have to learn how to write such a thing in rust and that it would be non trivial to do so, but that is just part of learning a new programming language to an expert level.
I think one misconception in here is that the rust "wouldn't look so pretty". You could do a line for line port of your c++ code using raw pointers and declare it safe, and the two implementations would have similar aesthetics. If your c++ implementation were memory safe to begin with, your rust implementation will be as well. The same process of thinking through the invariants and implementing them correctly is necessary in both languages.
I think the pushback you get that you find off-putting is against the sometimes expressed idea that it is a big knock against rust that you can't do linked data structures with zero cost abstractions in safe rust. This is why people say either "yeah but you can do minimal cost implementations in safe rust and zero cost implementations in unsafe rust". The pushback is against the idea that this isn't a good enough solution. But it totally is, it's certainly no worse than the situation in c++ where the division between a safe and unsafe subset of the language does not exist. Unsafe rust is no different than generic c++ code.
To your last point, there's a great book / reference called The Rustonomicon that is all about writing unsafe rust code well.
(I seem to recall a comment that it hasn't yet been fully updated for recent versions.)
"LRtDW is a series of articles putting Rust features in context for low-level C programmers [...] the sort of people who work on firmware, game engines, OS kernels, and the like. Basically, people like me."
Also, it seems like the author probably also wrote some low level libraries at Google... :)
I have heard claims that Rust would be performance-equivalent so many times and each time I bothered to check it turned out to be wrong. Especially in video game AI, indirection, pointers and linked lists are common. They are usually used in the performance-critical input-to-frame part of the game and there fitting things nicely into CPU cache lines is crucial for performance.
The end result is that I now have this vague feeling that Rust is a religious cult and I shouldn't take their claims at face value.
Unlike GC’d or dynamic languages, there is nothing intrinsic about Rust that makes it a fundamentally harder language to optimize. It’s really like a modern provably safe C++.
It also depends on how you code. If you make heavy use of functional constructs and complex types your Rust code will probably not end up being optimized as well. For tight algorithms it’s good to write “thin” Rust. The exact same is true for C++. You would not want to use a lot of STL or functional stuff in a rendering pipeline core. High performance C++ looks like C.
A lot of the optimizations required for fast code, like manual layout of custom datastructures are done by the programmers, not by the compiler. I don't doubt that rust can also express this optimizations given the similar low level control provided by the language, but what the OP and the parent are saying is that some of these are not idiomatic in rust and harder or impossible to express in the safe subset of the language.
> You would not want to use a lot of STL or functional stuff in a rendering pipeline core. High performance C++ looks like C.
FWIW that's not at all my experience.
In particular when it comes to unions which are heavily used in low-level code, Rust does optimizations that C and C++ can't even dream of (like all the niche optimizations).
In C++, sizeof(optional<T>) > sizeof(T) because the discriminat has to be stored somewhere.
This is true even if, e.g., you do something like `optional<T&>`. You know that T& is a non-null pointer, and you only have two variants, and one of the variants has no state, so you technically can encode this as 0x0 is the "no reference" variant, and the != 0x0 is the reference variant, and have `optional<T&>` have the same size as `T&`.
Rust does these layout optimizations of compressing the discriminant into gaps in the values of discriminated unions automatically.
So:
enum Option<T> {
Some(T),
None
}
for `Option<T&>` has the same size as `T&` in Rust, as opposed to C++.In C, an example would be:
struct DU {
enum { A, B } discriminant;
union {
bool A;
bool B;
}
};
You could encode that into 3 bits (i.e. have sizeof(DU) == 1), but instead you'll have at least sizeof(DU) == 2, because you need one byte for the discriminant, and one byte for the payload.It is very easy to create values with gaps in Rust, but C doesn't really support doing this.
Another optimization are alignment optimizations. In C, if you write:
struct S {
uint8_t a;
uint32_t b;
uint8_t c;
};
that ends up being 12 bytes long. In Rust, by default, that gets reordered as uint32_t, uint8_t, uint8_t, so it only ends up being 8 bytes.If you want that instead to be laid out like in C, you can write:
#[repr(C)]
struct S {
a: u8,
b: u32,
c: u8
}
and then you get the same 12 bytes as in C. There are many supported `repr(...)` options supported for algebraic data types, e.g., you can use repr(u32) for the Rust enum above to store the discriminant in a u32, and get the same layout as DU in C.But discriminant-less custom optionals classes with 'zero' type support via traits are easy to do (and I have done it many times). You do not have to rely on the compiler identifying a safe empty state and you can define any application specific one. For example if for one specific (and common) use case strings are always non null, optional<string, non_null_trait> has an obvious implementation.
So, yes, these sorts of optimizations are not done by the compiler has it has no notion of discriminated types (one day maybe...), but can be done generically by the programmer. In fact you could in principle optimize multiple layers of variant<optional<...>,... > and collapse everything in one discriminant,as long as you do not provide reference access to the sub variants; the required metaprogram is not going to be pretty though.
One of the advantages of C++ is that it allows this sort of control.
But I believe what the OP showed is that Rust also allows you to have that level of control. However, the default behavior for rust is to do the optimization that you have to go out of your way to manually implement in C++.
Rust isn't taking away control from what you can do in C++, instead, it's made the idiomatic approach one that is well optimized by the compiler.
I feel that this sort of belief relies too heavily on a presumption that performance is somehow proportional to the age of a programming language, and in the process ignores the fact that a) it's already using a highly optimized toolchain that benefits from decades of research and development, and b) there is absolutely no suggestion that any potential performance gains are relevant or exclusive to the Rust's front-end.
BTW Rust has some features that make optimization a lot easier, like the ability to assume strict aliasing in more cases. It’s safety features also make it easier to write zero copy, copy on write, and other optimization patterns without making mistakes. This enables some bold performance optimizations at the code level that would require serious bravery in C.
Long term I expect that Rust will be faster in some cases.
For linked list, you can use an unsafe implementation if that is performance critical. Unsafe implementation can be wrapped in a safe API for others to use, and this is how the standard library is implemented.
There is also a new paper about structures with internal sharing such as graphs without runtime overhead that Rc/RefCell would cause [2].
[1]: https://doc.rust-lang.org/reference/type-layout.html#the-c-r...
I had this feeling from the very beginning just by the tone of its zealots here and elsewhere. My theory is that since the language is very hard to learn but doesn't provide a lot of tangible benefits (memory safety exists in most mainstream languages fot decades), the sunk cost fallacy hits very hard. Then the only way to get some benefits for the time spent is to recruit new comers in a sort of pyramidal scheme structured around experience in the language. This way the zealots become priests and are compensated in form of social status. Of course, the whole scheme cramble if nobody joins, which is why the community is so aggressive (RIIR, vocal advocating, etc.) with the non-believers.
On the side of people coming from other memory safe languages, the benefit is much like Go, which has also become very widely used very quickly, so it seems worth taking seriously that the existing memory safe languages were maybe not satisfying people in some way. I think the popularity of both languages in this space can be boiled down to lowering the amount of abstraction without giving up memory safety, and simplifying the toolchain. Both Go and Rust do both of those things (in my view). I prefer rust because I don't like how the Go type system requires so much repetition of implementation due to lacking generics, but both languages make some amount of sense in this niche IMO.
Other people like rust because they came from C or C++ and got tired of some of their weaknesses but did not want to give up many of their strengths. Of course whether rust requires giving up the strengths of this languages is debatable and this where most of the backlash comes from. But rust is one of a very very few languages (maybe also D and Zig?) that can plausibly make this claim at all, and I think is has become the most actively developed and attracted the largest community of those, which does matter.
For me, it's a huge advantage to be able to plausibly use the same language for everything, from kernel drivers to giant applications. The only real competitor in that space is C++, and I think its cracks show more on both extremes, that is, I think rust is both a better C replacement at the very low end (see people like Linus Torvalds and Bryan Cantrill who are C diehards who dislike C++ but are interested or all in on rust) and is much better at the high level because of easy memory safety.
So I dunno, maybe there's a cult, but there are also very straightforward reasons why a lot of people are attracted to the language.
I'd say even if many of technical claims about Rust are true. It still seems very much like a cult.
Particularly in video games.
---
> The end result is that I now have this vague feeling that Rust is a religious cult and I shouldn't take their claims at face value.
Also, you realize that Rust allows writing efficient doubly-linked lists right ? And that the Rust standard library provides one for you right?
The whole argument of the OP is that people learning programming need to start doing so by learning how to write doubly linked lists because performance trumps all and therefore all children, elderly, and even university students that learn programming with javascript are lesser human beings and don't deserve our respect.
But yes, lack of goto and computed goto are deficiencies, not strengths.
The issue of course is that while these abstractions are zero-overhead and can sometimes be used safely, they aren't compositional in general. That is, they impose requirements on outside code which aren't easily captured by Rust's type system. This is exactly what the 'unsafe' facility in Rust was made for. Note also that more recently-developed abstractions of the "Qcell" or "GhostCell" type can in fact implement linked lists safely, and once these are better understood a variety of them will likely be included in the Rust standard library.
You can also say almost the same things you are saying here about C++. Type punning is almost always undefined behavior. There is no way to construct an array of variable size at an address that you specify without a pointer indirection. But with ASM you've got no such problems!
That's probably the only reason to use them. If you don't care about performance there are simpler data-structures available.
If you need to implement an algorithm for which you need O(1) splice, then doubly-linked lists are a data-structure that give you that. If your objects are very big the cache misses might not matter that much, and neither would ref counting.
If you go one step higher, and can modify your object data types, and are careful with how you allocate your data, then intrusive doubly-linked lists can give you equivalent performance to a vector with better algorithmic complexity for many insertion / removal / splice operations, etc.
The stars do however need to align a lot for a non-intrusive doubly-linked list, like the one being discussed above, to be the best answer for whatever performance / algorithmic problem you are having.
That O(1) splice must also be accompanied with an iteration. Which, in my experience, is really rare. If that iteration step wasn't already a part of the splice requirement then it can often be faster to do the splice via a memcopy.
My favorite algorithmic mistake was someone at my company used a binary search to maintain sort order on a linked list. IIRC, that turns insertion into something like an O(n^(log n)) operation whereas it's O(log n) operation on an array.
I don't think it necessarily must, but it tends to be.
People tend to keep pointers to elements of the list all over the place, so if you have the right 3 pointers (being and end of list you want to insert, and position in another list), then you can do it in O(1).
If not, and you need to traverse the lists... as you mentioned there are other data-structures that might be much better.
No, you aren't.
None of these languages catch data-races, so writing multi-threaded code is pretty much as hard and error prone as in C. Other higher-level languages like Python fix this by holding a global mutex, so that using multiple threads doesn't really buy you that much there since threads always execute sequentially to avoid data-races.
This is in contrast to Rust, which allows anybody, even people without experience in low-level programming to accelerate their applications using multiple threads, without introducing bugs.
Actually, Mozilla tried to multi-thread Firefox multiple times using C++, and failed, over and over again, because every attempt would introduce subtle bugs. Rust was created to address this issue.
All real world languages have facilities for various kinds of safe concurrency, e.g. actors and other kinds of message passing. When it comes to concurrency, Rust isn't anything special. It's not even that good.
Rust has one killer feature: memory safety without GC. Except for this feature, Rust is mediocre. If you don't need the memory safety without GC, you can use almost anything else and be better off.
In Rust, many libraries especially the standard one have lots of unsafe code, with a history of security bugs there [1]. In C#, the standard library is written in almost 100% safe code.
About threading, while it doesn't catch all data races automatically, C# implements many useful things on the VM level. Monitor class, or memory model guarantees, are very hard to implement in languages who compile to native code.
[1] https://shnatsel.medium.com/how-rusts-standard-library-was-v...
Citation needed: which part of my comment says this?
Also note that the author of the article is the actual inventor of RCU.
I know, so?
> On the other hand the author is probably saying that the kernel RCU C API itself has been proven correct,
But this is not true right?
There is no formal proof, e.g., in Coq, that proves a useful subset of the RCU C API correct (as in, if you stick to this useful subset, you can't introduce UB).
OTOH, a safe Rust API could be proved correct to use at least, and an implementation of it that uses unsafe, like the one in crossbeam-rs, is typically very amenable to theorem provers (much more than C).
So what I am saying is that the author is ignoring the value of providing a safe Rust API for RCU, like crossbeam-rs.
If the Linux RCU APIs can't be mapped to it, then maybe they are incorrect to use, and this is why we can't prove them correct.
Doing the work of mapping them, and understanding any issues, could lead to changes to those APIs that might allow us to prove them correct.
Of course that's still not proving the actual C code, the translation probably still need to checked by hand.
Just wondering, how many unsafe code in commonly used libraries (standard library, tokio etc.) are proved using theorem provers?
Some of the unsafe code in the standard library also.
Some parts of crossbeam have pencil&paper proofs documented.
There's a whole proven correct kernel: http://sel4.systems/
So yes, there is a correctness proof for a completely different operating system that does not use the one thing we are talking about here.
Adding an unsafe block around a call to C code doesn't prove anything. Rust does not provide proofs of the code you write in it, except if you do not use unsafe at all.
RCU is an API and a general technique. It can be implemented a ton of different ways (including, as the author noted, as a single instruction if Linux is built with pre-emption disabled). You must not conflate the the soundness of an API with the soundness of its implementation. 'Can Rust prove Linux RCU APIs correct?' is nearly completely without meaning.
A similarly meaningless project would be to prove that the idea of malloc/free is correct. There's nothing that could even be incorrect or correct about it, and you don't need Rust wrappers to tell you that. Now prove glibc's malloc is correct; very different question.
I didn't say otherwise.
What I said is that writing unsafe code is writing a proof that the code is correct.
If the proof is incorrect, the behavior is undefined, and Rust makes no guarantees.
unsafe literally means "I've proven this code correct".
As mentioned, most uses of unsafe in the standard library are accompanied by a comment that documents the proof.
The narrower point I think you're making is that hopefully attempting to create a safe API around the RCU pattern makes someone think about RCU's correctness even harder. I'm not sure it will produce much, given how much attention RCU has already received in the past 20 years. Any innovation you get from that process will probably be concentrated in describing/encoding the ownership pattern in Rust terms, since that's non-obvious. But the epoch GC implemented in Crossbeam is pretty similar overall to RCU, so it has much to build on in that regard.
I disagree here (and that's ok).
The unsafe keyword tells Rust "this code is sound", i.e., there exist no program execution for which this code will cause safe Rust to exhibit undefined behavior.
Rust will parse it, type check it, compile it, link it, etc. "as-if" it was sound.
And Rust only makes any guarantees about what your program does if the unsafe code was actually sound.
From the point of view of the compiler, the unsafe block itself is a user-provided proof that whatever that code does is sound.
If you write: unsafe { *ptr } that's a proof that dereferencing the pointer is sound, therefore the pointer is non-null, it points to an aligned allocation able to hold a type, the value at that memory location is a value of that type (i.e. the pointer is dereferenceable), etc.
The Rust compiler accepts your proof and generates code as if your proof is correct. In particular, it will compile code around that code under that assumption, it will eliminate null pointer checks after that code, etc.
If the proof that an unsafe block provides is incorrect, then the behavior is undefined, and Rust makes no guarantees.
I agree with you that given some unsafe code, actually proving it correct is not something that the compiler does, and that creating a tool that does this is hard. If it were easy, the compiler would already do it, and we wouldn't need unsafe at all.
Unsafe exists because there are things that the compiler can't prove correct, and this escape hatch allows the user to provide a proof that the compiler can integrate into how it compiles the program.
There is a whole repository of Iris proofs for unsafe blocks in the Rust standard library, for which we have formal proofs. But these are not integrated in any way with the Rust code one writes. They have been written by humans aiming to get a confirmation that the proof expressed by the unsafe blocks are indeed correct. Some were not, and had to be fixed, and pretty much every fix has generated a paper.
Yes. Unsafe blocks are bare assertions, no more. Everything you described about what happens in Rustc when you use it is better described as making a bare assertion to the compiler. Everything beyond that is socially constructed. Saying it's a proof sounds even more silly when you recall that Rust programmers treat it as quite the opposite -- radioactive markers of places where the program might cause UB.
You were trying to characterise it as a proof because you're claiming that "Can Rust prove Linux RCU APIs correct?" is an interesting question. But you are actually asking "Will Rust compile RCU code if we assert to it that using the RCU API is safe?", and the answer is unconditionally yes. It is not an interesting question at all. Your misuse of the word proof led you there, so I am suggesting you stop calling the unsafe keyword a proof.
In the broader scheme of things, yes, the way in which Rust encourages you to isolate unsafety is helpful, and so if you wanted something formally verified Rust is a good language for that. That doesn't really translate to RCU, because the hypothetical Rust API would not be an implementation of RCU. All the Iris proofs that have come out of the RustBelt project are bad examples here, because they were actual implementations of things, not APIs around black box implementations that change under your feet when you set kconfigs.
As a final thought, here is a supposedly safe-encapsulated Rust API for RCU, I would invite anyone to suggest how writing down this very simple API or similar + a few calls to extern "C" wrappers for the RCU macros could possibly help "prove Linux RCU APIs correct", whatever that means: https://docs.rs/rcu-clean/0.1.6/rcu_clean/struct.ArcRcu.html
(Little edit: if you're interested in what circumstances will lead Rust wrappers of C code to elucidate safety, look into rlua, a project to create a safe wrapper for Lua (https://github.com/amethyst/rlua). Arguably the main outcome is that we now know that the Lua API cannot be wrapped safely with zero cost, in no small part due to the use of longjmp. Some of RCU's APIs require you do not sleep or block in any way while using it, lest you wake up with the GC train having left the station. For this reason I imagine you won't be able to encode all the usage requirements in the Rust type system, let alone make an actual safe zero-cost API, let alone turn that into a proof of anything.)
(In fact, there was much bikeshedding at one point about renaming the 'unsafe' keyword to make this explicit: suggestioned included 'assumed_safe', 'trustme', 'trusted', as well as less serious ones like 'yolo' and 'hold_my_beer_and_watch_this': https://github.com/rust-lang/rfcs/pull/117)
I have corrected many math exams of students over the years.
A "significant" amount of the "proofs" provided by the students in those exams were wrong in one way or another.
Many students continued the exams under the assumption that what they just incorrectly proved, was actually correct.
They had no way to verify whether they proofs were correct or not. No theorem provers were allowed in the exam.
Yet all of them called what they wrote "a proof".
---
This is pretty similar to how Rust works.
Without unsafe, Rust proves your programs sound for you.
Unsafe let's you supply your own soundness proofs for parts of the program that the compiler does not know how to prove (or disprove).
If the proofs are not correct, RUst makes no guarantees.
The claim that because "Rust does not prove these for you, they are not proofs" is unfair.
That's like claiming that any "proof" that has not been proved in a theorem prover does not deserve to be called "a proof".
I respect your way to see this. I just don't see it this way.
If you are aware of one, I am intrinsically interested in this topic.
Memory safety without GC is a big selling point of Rust. However, sometimes your program’s ownership model does not fit with what the Rust borrow-checker can verify. Not to worry, just use “unsafe”. But if “unsafe” is used, can we still claim that Rust code is safe?
One might respond that you can do the same kind of modularization in C++. However, I think a Rust advocate would argue that it's much more difficult in practice to modularize unsafe code in C++ to the same degree, since unsafe idioms are deeply embedded in the language and library ecosystem.
My experience is that you hardly ever need unsafe Rust code. If your ownership model doesn't the borrow checker, there are of course also other avenues like ref counts or even an actual garbage collecter (which you can chose to apply only to values whose ownership situation is not aligned with the stack). Only when those options are unable to give you the performance you want you'll need to reach for unsafe code.
Other scenarios where you'll need unsafe might include things like SIMD (although safe wrappers exist) or FFI, but hopefully we'll evolve the Rust ecosystem over time to concentrate the unsafe code in a few places by building abstractions, where they are thoroughly reviewed and well tested.
I doubt that experience applies equally to kernel development.
Of course, the kernel:: -modules it uses undoubtedly have more, but I imagine the idea is to provide abstractions that allow actual driver code not to use unsafe.
And then there's stuff like file system code, that—I imagine—would require no more uses of unsafe than a database server written in Rust.
For example suppose that we're back in the late 1990s and so for some reason dozens of different yet roughly equivalent 100Mbps fast Ethernet PCI cards are available, each offering the same functionality but with slightly different register layouts, interrupts and quirks. We are of course writing drivers for them in Rust.
Obviously at some fundamental level "configure PCI bus" isn't safe - if you get it wrong that'll be bad. But as the author of mediocre Ethernet driver #42 you aren't implementing "configure PCI bus" you're just calling a safe abstraction that worked for the other 41 drivers.
Handling userspace calls to twiddle the Ethernet interface is potentially unsafe too, but likewise that will have been factored out into common code and you're just talking to a safe abstraction.
Rust's type system and borrow checking makes it easier to write a safe abstraction for relatively low level ideas like "Only one device can own this MMIO register" or even "this can be 0x03 or 0x06 or even 0x18 but it can't be all zeroes, that's not a thing".
This sort of stuff is boring and in principle already not particularly unsafe in C today, this sort of abstraction is already done - but of course in C you've no practical way to get the promises Rust offers for safe code.
Just as code written in C might be safe, it just can't be proved to be safe by the compiler! Rust's advantage over C is that the machine-unprovability is confined to small, defined zones, where hopefully human programmers can, formally or informally, prove the code safe.
Because compiler cannot prove that what you're doing is safe it is your responsibility to do so. Once you've _proven_ that your unsafe code is safe, then it cannot have propagating effects in the rest of the code.
A lot of people get this confused. They think unsafe means you can do whatever you want. In reality it means you'll play by the rules even if the compiler cannot verify them all. In truth it is not as hard as it sounds, we do it in C all the time.
`unsafe` in rust really means that some of the compiler guardrails are disabled:
* you can call unsafe (and extern) functions
* you can dereference raw pointers (this has significant consequences such as being able to create mutable references out of thin air)
* you can access or modify mutable statics
* you can access `union` contents
It is very easy, particularly coming from another language, to butt heads with the language's design and then try to solve it with "trust_me" blocks. Throwing everything in unsafe feels like you're doing it wrong, and correctly so.
The downside is: Every time this discussion comes up, someone has to explain what "unsafe" actually means. Every single tech talk ever that touches on "unsafe", pre-fixes the talk with this explanation too.
A more funny idea instead of "trust_me" would be "safe_because[...]" where "..." is a rationale about the code plus your social security number and biometrics data.
For any control freaks and regulators out there: If you take the above seriously and try to implement such a thing you will be burnt by the stake by cyber elves.
You want to do something that the compiler can't understand but you know how it works? Ok, write down _why_ it works so others can know that, too.
Would it be possible to write the above as a Rust macro, and mark invocations of the ordinary "unsafe" operator as a warning / linter alert?
PS: I actually wrapped most F# unsafe (= partial) functions in our codebase in such a way. `Option.get x`, which extracts a possibly-null value and crasheswise, is deprecated and replaced by `Option.getOrFail fmt args x`, where you say why you are sure the value can't be null and provide context variables to diagnose a possible failure.
e.g.
let title =
textLines
|> List.headOrFail "We generated the line lists by subgrouping the original list %A, so we know they're non-empty" originalList[1]: https://rust-lang.github.io/rust-clippy/master/#missing_safe...
But the keyword "unsafe" has been used in C# for quite a while, and probably in other languages too. So by the time Rust came out, it was already a well understood concept.
In the C# manual: "the execution engine works to ensure that unsafe code cannot be executed in an untrusted environment". Still that notion of trust.
One problem I see with "trust_me" or maybe "trusted" is that it gives the opposite message from a user point of view. Trusted code is code the compiler should trust but that you should not trust.
The main issue is that it came late in D lifetime, so until situation improves, @system is the default, because there's too much std/legacy code which still assumes @system as default. Many new projects however start with @safe as default.
This tends to be a beginner issue for c/c++ programmers. 99% of the time you just need to rethink your design into something better.
- limit the size and scope of the unsafe code
- make the unsafe part stand out, "here be dragons", so you know when to be particularly alert
- act as a safe way to encapsulate unsafe code so that it can be used by safe code, and used without the lack of safety propagating outside of the unsafe zone
By the way, several parts of the std (Mutex and Arc for example) are written in unsafe Rust. They have to be, by necessity, because they allow for aliasing and mutation.
C/C++ is not just "everything is unsafe", it's also lacking many powerful features that exist even inside unsafe{} blocks: ergonomic enums with associated data, type-checked generics, ergonomic moves etc.
Unsafe code is not the only solution to borrow checking problems, as Rust supports reference counting, which is an implementation of shared ownership.
Whether this defeats the point of Rust is subjective, but regardless, it's a solution to at least certain ownership problems.
A very typical area where you can observe this is bidirectional trees. They're ugly in Rust! :)
Now, memory unsafety is not a black and white matter. In my experience with the community, you'll be surprised at how much pragmatism there is. Nobody will scream sacrilege if one expresses the necessity to use memory unsafe code¹.
The connection to your question is that overally, the Rust memory safety model is not all-or-nothing, rather, a spectrum of safety above the convenientional entirely-unsafe approach. With this in mind, I'd say that "98% Rust-safety" is well beyond the "100% unsafe" approach of traditional low-level languages.
As a matter of fact, I suppose that in the future, we'll see languages developing intermediate memory safety features (I think Zig was doing something like that).
[¹] The actix-web case is a very interesting one. I haven't looked at it, but my personal suspicion is that it made use of _wildly_ unsafe code, and that usafe of unsafe code was not the problem, rather that unnecessary usage was.
IIRC it was mostly the initial reluctance of the author to merge PRs that rewrote unsafe to safe Rust. Some escalation, some emotions, and then everything moved forward in harmony and... with more safety.
No, not really. Spinning the actix-web as the author's reluctance to accept PRs is a gross misrepresentation of the problem faced by the project.
The author even went to great strides to explain that his use of unsafe code was indeed safe.
For context, here's what the author had to say about that regrettable episode.
No, not really. The author, a professional at Microsoft, was a hobbyist with rust who (by his confession) wished to explore the limits of performance using various techniques. People made a mountain out of the fact that his project rocketed up to the top of the TechEmpower benchmarks -- as if this obligated him to conform to their flash-mob wishes regarding sanitized code, whereas the project very explicitly made it clear that production use was at one's own risk.
What happened to him was a hit-job by a bunch of entitled bullies. No less than Steve Klabnik acknowleged this ("far, far, far over the line"[0]). The fact that ntex exists today[1] yet the world has somehow not exploded yet demonstrates how overblown the criticisms were.
It's very hard to believe that "a hobbyist" would, by mere chance, write the absolute best performing framework in the world.
https://www.techempower.com/benchmarks/
Your comment not only lacks credibility but also sounds extremely petty.
And that was especially true of the benchmark code itself.
TechEmpower's methodology, additionally, is not perfect. IIRC they support HTTP 1.0 only, which is not what most modern servers are focusing on, but actix-web did.
There was some pretty bad misuse of unsafe in actix-web. And as I recall it, the author was resistant to fixing it, even when patches were provided. This is the author's right. It's their project. But it's anyone else's right to make this sort of thing visible, and to advocate against holding up actix-web (at the time) as a good example of what Rust programmers could deliver.
A lot of people can't or won't communicate effectively, and often blur these lines. Either assuming there is entitlement where there is none, or expressing entitlement when they didn't mean to. But there is a reasonable position underlying what happened, and unfortunately some took it too far.
Some places of code were actually unsound (which means: got function containing unsafe and marked safe even if it didn't upheld the invariant of the unsafe block, meaning you could trigger UB from safe code: this is a big no-go in Rust).
That being said, the whole drama with people mobbing the author on every social media, calling him name and so on, is indeed regrettable.
That's what any author of unsafe code is supposed to do right? Explain why the unsafe code is in fact safe (enough).
There's a common trope among people unfamiliar with rust where they assume that if you use unsafe at all, then it's just as unsafe as C and rust provided no benefit. Comparing C's approach to safety vs Rust's is like comparing an open world assumption to a closed world assumption in formal logic systems. In C, you publish your api if it's possible to use correctly (open world). In Rust, you publish a safe api if it's im possible to use in correctly (closed world). Rust's key innovation here is that it enables you to build a 'bridge' from open world (unsafe) to a closed world (safe), a seemingly impossible feat that feels like somehow pairwise reducing an uncountable infinity with a countable infinity. Rust's decision to design an analogous closed-world assumption for safe code is extremely powerful, but it seems very hard for old school C programmers to wrap their head around it.
My point being: Unsafe does not invalidate the safety of the overall code, it shows you places where you have to be extra carefully. On the other hand if you yourself do not use unsafe and all the libraries you use are of high quality you can be sure your code will be correct. Which is the best you can do in reality.
Keep in mind that GC and a large parts of low-level code that is fundamental for languages that have GC is written in C, and we still are findings safety bugs there.
Writing unsafe requires a certain degree of skill, because one needs a deep understanding about what guarantees safe Rust provides, to ensure that your abstraction does not break them.
If you read the standard library docs, every unsafe block has a comment explaining / proving why it can't cause safe Rust code to exhibit undefined behavior.
Many of these blocks, particularly those in libcore, have been proved formally correct in proof assistants.
There is a theorem that proves that "if a safe Rust wrapper over unsafe code is proven correct, then extending safe Rust with that wrapper is still sound".
So you basically can infinitely extend the safe Rust subset of Rust with these abstractions.
> But if “unsafe” is used, can we still claim that Rust code is safe?
So the answer to this is "no, you can't claim that, you have to actually go and prove it".
Often, very often actually, the proofs are trivial.
For example, to index into a slice:
fn index_slice(slice: &[T], i: usize)
// SAFETY: if the index is not in bounds, we panic
assert!(i < slice.len());
unsafe {
*slice.as_ptr().add(i)
}
}
suffices.Why? Because code that creates a slice with a ptr and a len field that do not point to a valid allocation with len valid elements already exhibits UB. That is, for any valid Rust program, we are guaranteed here that these two fields are "ok".
So the only thing we need to make sure of is that the index "i" is inbounds, and that's trivial to do with an assert that panics if this is not the case. That is, at runtime, no program for which the index is not inbounds will reach the ptr.add(i) method, so there is no way to offset this pointer out of bounds, much less dereference it.
Basically, because all safe Rust code can rely on all other safe Rust code upholding the rules, most of the proofs are really easy. The fact that you can prove safe abstractions over unsafe code in isolation from other abstractions is one of the main features of Rust.
---
For synchronization primitives, proving the absence of data-races is particularly hard, and requires a lot of expertise. The subset of Rust programmers that actually need to do this is very small. Most people just use those primitives that have already been proven today, of which there are many.
My take is that Rust is proposed based on the value proposition of its Safe Rust feature, but in the Linux kernel that feature has limited use.
Therefore, if Safe Rust is out of the table then what's the point of writing/rewriting parts of the Linux kernel in Rust? What's the value proposition?
Where does your take come from?
Everyone I've heard of proposing Rust for the Linux kernel does so almost exclusively by making the argument that Rust can create safe performant abstractions over unsafe code.
That is, instead of requiring all Linux kernels programmers to be deeply familiar with RCU, you can have a small group of RCU experts create a safe abstraction over RCU, that then all other Linux kernel programmers can just use, without having to understand very deeply how it works. If they make a mistake, their code won't compile, and the error message explains why and how to fix it.
This is a huge productivity boost for large complex projects, and a huge safety boost since these abstractions prevent a large set of security vulnerabilities, and these are the main reason I've seen people to use to introduce Rust into the Linux kernel.
The fact that Rust does this with minimal - often zero - runtime overhead is a requirement for the kernel. But even for abstractions for which this is not the case, it still allows people to only have to become an expert when they want to improve performance, and any use of unsafe immediately triggers a requirement for such code to be reviewed by the actual experts.
This is also a huge boon for these experts, since it substantially reduces the code they have to review. E.g. now they only have to review uses of the "unsafe RCU" APIs, instead of also having to review the uses of the safe RCU APIs, which is the case today.
And also a big boon typically for the project in general, since it is much easier to attract "newbies" if there is stuff they can break. A newbie hacking on the linux kernel 20 years ago had much much much less to learn to become proficient than a newbie today. Reducing what newbies have to learn is a surprisingly effective way to increasing the contributor base of a project.
From this very discussion, for starters.
> Everyone I've heard of proposing Rust for the Linux kernel does so almost exclusively by making the argument that Rust can create safe performant abstractions over unsafe code.
The whole point is that the same thing (creating safe abstractions over unsafe code) is what developers have been doing with C for close to half a century.
Therefore, there must be a value proposition regarding the use of Rust other than Safe Rust. I haven't heard one so far. In fact, I've read the exact opposite, such as compilers needing rewrites to accommodate that usecase.
Thus, with Safe Rust out of the table, what exactly is Rust's value proposition?
> This is a huge productivity boost for large complex projects, and a huge safety boost since these abstractions prevent a large set of security vulnerabilities, and these are the main reason I've seen people to use to introduce Rust into the Linux kernel.
I feel those baseless assertions require some supporting evidence. Is there actually any tangible and concrete evidence that Rust, specially Unsafe Rust, improves productivity and safety? I'd love to read about it.
They have been trying to do this, but it doesn't work because C's support for this is very poor. Not even C++ managed to do this.
Rust makes it easy to do this in a reliable and efficient way that is 100% guaranteed to work.
> I feel those baseless assertions require some supporting evidence. Is there actually any tangible and concrete evidence that Rust, specially Unsafe Rust, improves productivity and safety?
Yes, see: https://prev.rust-lang.org/en-US/whitepapers.html and https://www.rust-lang.org/production , and well https://dl.acm.org/action/doSearch?AllField=Rust (there are a bunch of peer-reviewed papers there that show this, serve yourself).
It is more like providing an "axiom" as the compiler won't be able to check it and instead has to assume that it true.
I wonder if there is scope to extend the language in the future to actually add machine checked proofs.
Normally full machine checking doesn't scale to large projects, but if one had to restrict to writing proofs only for the small subset of unsafe code and let the simpler type/lifetime system handle the rest of the code it might be manageable.
I don't think one should see this as an axiom, although as you mention one could.
If your unsafe code allows safe Rust to exhibit undefined behavior, Rust makes no guarantees whatsoever about what your program will do.
The convention in the standard library is to provide an actual proof that this won't happen as a comment to each unsafe block.
Right now, there is no tooling to check these proofs, but there is a lot of ongoing research about how to do that, and the standard library already has some tooling to check things about this.
> Normally full machine checking doesn't scale to large projects, but if one had to restrict to writing proofs only for the small subset of unsafe code and let the simpler type/lifetime system handle the rest of the code it might be manageable.
That's the idea.
It is not as simple as just proving each unsafe block independently, but rather all unsafe blocks within a Rust module must be proven together.
From this point-of-view, it would make sense to only implement one unsafe abstraction per module, and keep these modules as small as possible, providing most of the implementation of the abstraction outside the module using only the safe API.
Turns out, this is already how Rust code is written, since this is not only useful for writing proofs (either via assistants or using pencil and paper), but also helps humans reason about the correctness of the unsafe code by keeping all the "fluff" away.
Now, beyond obvious things a beginner would want, you start to get into some pretty cool structures, and many of those are not provided by Rust's standard library, however, if you aren't excited about implementing them yourself others have likely done so. For example maybe you've read about "Hazard pointers" for very concurrent systems where you can't afford to take a lock when destroying shared resources but you want to only destroy resources that are no longer used. Rust doesn't provide Hazard pointers in its standard library, but several crates offer different takes on this.
Again, the implementations are unsafe internally, but if you trust that the programmer got them right, your use of their library is safe.
You see it used every so often.
For these reasons, most garbage collection systems are using mark and sweep, or something similar.
That links to a message saying this:
> "ksmbd was then merged unbeknownst to me and my worries confirmed to be true as some out-of-bounds bugs (which are impossible in Rust) surfaced immediately."
Sad.
The only bugs I found have a long email chain with people being surprised that the kernel module even passed any tests at all, with extensive changes needed to fix logic issues for the path conversion code. The devs. even added an unload reload cycle to the tests because the module was in a bad enough state that it wouldn't consistently run over several iterations. I wouldn't be surprised if they would simply turn to unsafe if they had to write it in Rust, errors are apparently to be worked around not fixed. [1]
[1]https://lore.kernel.org/lkml/CAH2r5mvu5wTcgoR-EeXLcoZOvhEiMR...
edit: the last part of my comment may be a bit over the top, but I think the quality issues at the language level are just the tip of the iceberg for this module.
They probably would. But it's easier to find usages of unsafe through grep or the like than it is with C because Rust's syntax and requirement of marking unsafe make it easier to find. To use one of my own codebases for example, I can look at the instances of the following regex patterns:
* "\bfn\b" (function declarations), 138 matches.
* "\bunsafe\s+fn\b" or "\bunsafe\s+extern\b" (unsafe function declarations), 23 total.
* "\bunsafe\s*\{" (opening an unsafe block), 41 matches.
* "\bunsafe\s+impl" (unsafe trait implementation), 0 matches.
We can also look at the total lines of code written (2269), and ask if this these numbers feel reasonable for what it's doing. Obviously that will require familiarity with Rust, what's being implemented, and the expected amount of unsafe based on previous usage in the kernel, but it would at least give an overall feel. And, of course, if you want to examine any of them, you know exactly where they are.Is that what happened?
For drivers it could be fine. But so could many other languages. So why Rust?
$(language of the week) changes, and people continue using the old software, written in c/c++/...
torvalds said recently: “Probably next year, we’ll start seeing some first intrepid modules being written in Rust, and maybe being integrated in the mainline kernel.” [1]
the rust community is a bit over-excited, but if you can look past that, there's substance in the hype.
[1] https://thenewstack.io/linus-torvalds-on-community-rust-and-...
Rust is 10 years old. See, such knee-jerk reactions are exactly why many people have trouble taking the anti-Rust arguers seriously. It's like you people defend religion or something. How about we stick to facts?
C is a very complex language as well. It's just that when you're not aware of all the complexities in Rust, the compiler will refuse to compile; in C, you'll get a security vulnerability or weird crashes at runtime. Yes, this is a generalization, but only slightly.
C is deceptively simple – its complexities are well hidden.
Rust's corresponding layer is much thicker due to its memory requirements and historical origins.
No, it doesn't, and it hasn't for quite some time – not since C compilers started exploiting undefined behavior for optimization opportunities. It used to be that if you wrote
a + b
where a and b are int values, the compiler would emit an ADD instruction in assembly. This isn't necessarily true anymore, because compilers are allowed to assume that signed overflow doesn't happen, which makes writing correct overflow-detection so difficult and non-intuitive. Even worse, code which was working correctly for 20 years might start exhibiting subtle bugs from one day to the next, just because you've upgraded your compiler from version 9.4.2 to 9.4.3.Undefined behavior is one major reason why writing correct C code is complex and difficult in practice.
it's the exact opposite - if you want your compiler to just emit an ADD no matter which platform you are on, then it cannot be the language that defines the semantics of "+" in the corner cases (thus UB). If you define it as overflow on 2's complement, then on a DSP with saturating arithmetic your compiler will have to emit additional instructions to detect if you are saturating and wrap instead.
> 145 pages
Yes, and (within that 145) they take a whole page to define the word "scope". Not even in the context of C!! Just the general language-agnostic term "Scope". That's hardly meaningful...
But, does that in turn make C more complex? It definitely makes it less safe. But I don't know if it automatically makes it more complex.
If it weren't hidden, it would be a different language. Would that language be less complex? That depends on what that language would look like.
> But, does that in turn make C more complex?
More complex than what? I'm not saying that C is more or less complex than Rust – what I'm saying is that C is not simple – its considerable complexity isn't visible at first glance, but it's there.
Sounds like a total hand-wave to me. Can you provide examples of "well hidden" complexities?
Be honest with yourself: you NEVER made a mistake that made you facepalm sometime in the future? Not once inn your life?
People just refuse to believe they can make mistakes. That's why those discussions are impossible to win; they are about personal maturity and not about programming prowess. On one side you have the people confidently saying their famous last words: "it's simple, you all dumb or something?" and on the other you have those that are like "I know I can screw up even if I'm the best ever so let me use a language that will catch part of my mistakes".
But yeah, the latter group can never convince the former. You all just repeat the same things every time and always miss the forest for a single tree. Very frustrating.
Ew
If something can be done from userspace, it should be done from userspace.
Legacy Linux drivers could be sandboxed in user space inside a fake Linux kernel environment, a bit like how captive NDIS used to provide a fake NT kernel for network drivers run under Linux.
The fact that the kernel ABI is unstable is irrelevant because you're not relying on ABI stability. You're using a specific version of the core kennel paired with a specific version of the driver.
Yeah your kernel could totally support USB coffee makers using USB passthrough to a Linux VM as long as (1) you write a Linux daemon for each device that lets you talk to the hardware over RPCs from the host and (2) all application software that wants to talk to said hardware uses your dbus API. That's a hell of an ask, but it's doable.
But what about the xHCI drivers for USB? You still have to write that for every USB chip on the market.
What about file systems? I've been burned by FUSE performance.
What about GPU drivers? This sucks on Linux and it'll suck even harder on your hypothetical kernel.
Or any device that sits on the PCIe bus for that matter. There are so many drivers that rely on running in a shared address space with data structures and hardware state shared with the rest of the kernel. At the very least you'll have to support the 2 new motherboard chipsets that come out every year. To be clear, I'm just taking about desktops. Laptops and phones will be a whole new world of pain
> For all of these references, I give a big "Thank You!!!" to my co-authors.
...
> This pointer is now a "zombie pointer" that has come back from the dead, or that has at least come to have a more entertaining form of invalidity.
...
> Furthermore, even given unanticipated universal acclamation of Rust within the Linux kernel community combined with equally unanticipated advances in C-to-Rust translation capabilities, a significant fraction of the existing tens of millions of lines of Linux-kernel C code will persist for some time to come.
and his parting words, not that he is negative or pessimistic about Rust but rather highlighting some unsolved challenges and their possible solutions:
> In short, choose wisely and be very careful what you wish for! ;-)
If you take an average chunk of kernel code, and try to reason about how to rewrite it in Rust, you'll note that the result looks surprisingly like C.
Working with hardware registers, instructions, memory mapping, and trying to do all that with optimal runtime performance puts a lot of constraints on you.
Rust is all about memory safety. Well, what does it mean when the entire memory map is not only accessible to you, but DEFINED by you?
Language shapes how we think about the world. Ideas that are hard to describe in the language of the mind are hard to imagine in the first place. What potential avenues for improvement might we lose by shifting computing to safe but constricting languages like Rust?
Yes, it is, and the linked blog posts talk about how. Did you read them?
My point isn't that some ideas can't be expressed in Rust. It's true that with sufficient effort, Rust can express anything.
My point is that some ideas are harder and more awkward to express in Rust compared to C, and that as a result, people might have been less likely to invent those difficult-to-imagine ideas if they'd been working in Rust and not C.
I'm basically arguing for the "weak" variant of the Sapir-Whorf hypothesis as applied to programming languages. The weak form of Sapir-Whorf is obviously true with respect to human languages.
To the responses below:
Yes, it's true that Rust's higher level of abstraction allows for optimizations difficult to express in C. It's worth pointing out that C++ has similar expressiveness but doesn't constrain access to the machine in the same way.
> Sure some things might be harder to express in Rust, but the tradeoff is well worth it.
You're missing my point. For some types of code, the safety Rust provides might be "worth it". I'm just suggesting that there really are performance-enhancing techniques that would have gone uninvented in a Rust-only world.
You seem to be agreeing with me and adding on top that we don't need that stinking performance anyway, so nothing I said presents a problem. Can you see how someone might disagree with that perspective?
All of these are possible in an elegant/comfortable to use way in “higher-level” languages, like Rust and C++.
Ok, let's assume it is confining (something that I disagree with). So what? Seat belts are confining as well.
Sure some things might be harder to express in Rust, but the tradeoff is well worth it.
I'd rather take the confinement or Rust, rather than pervasive paranoia of whether I'm going to hit C UB anywhere.
[1] http://dtrace.org/blogs/bmc/2018/09/28/the-relative-performa...
The weak form of Sapir-Whorf is obviously true with respect to human languages.
No, it's not. Nothing is "obviously true", least of all a pretty grandiose theory, and I can't find any source that empirically backs up the idea that the weak version of the hypothesis is true at all.[1]: http://dtrace.org/blogs/bmc/2018/09/28/the-relative-performa...
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Just as safe and fast as C++, but with more work done to micro optimize.
They weren't trying to game benchmarks in their analysis.
Surpassed?
Looks kind-of similar.
It's a shame that a high-quality microkernel OS like QNX has never been made, though, except as proprietary commercial products. QNX has by most accounts been quite successful (the , but it's still a niche product.
I used to tell the QNX sales rep, "You worry about your product being pirated. Worry more about it being ignored." The major customers are big industrial firms and auto companies, and if they use it, you can find them. Besides, they want the support contract.
What a strong and damning claim, it would be a shame if it were left entirely unsubstantiated...
Example: see the drama surrounding https://news.ycombinator.com/item?id=28513130
The Linux kernel people are completely different. They're direct, aggressive, and have arguments directly, out in the open. Linus will call you an idiot in public, which is a big no-no in "nice" places like Rust, but if you fix whatever it is that's making him call you an idiot, Linus will then take your code and think nothing more about the incident.
By contrast, in Rust-land, disputes might simmer forever because nobody is allowed to by "mean" and bring them out into the open. Things will just mysteriously not happen for you and you'll get more and more frustrated that nothing seems to be making any sense.
In other words: Linux has an old-school hacker culture, and Rust has the culture of a college campus.
People used to the Rust style think of the Linux community as a bunch of toxic aggressive assholes. People used to the Linux style think of the Rust people as a bunch of toxic passive-aggressive assholes.
(I personally much prefer the Linux style.)
That was drama, yes. I didn't see any intrigue or backstabbing or superficial politeness.
By superficial politeness do you mean keeping certain things private?
It's mostly sheltered Americans who misinterpret Linus's way of doing things as "toxic". Said sheltered Americans also coasted their way through life without ever having a Linus-esque reality check, so again they perceive that kind of concise communication as a personal material attack on them. Is Linus a bit too curt? Maybe, but he gets things done.
Super strange comment, your one. Goes into "Rust devs can't take criticism" land and has no leg to stand on at that.