What Is Rust's Unsafe? (2019)
nora.codes
nora.codes
One use of "unsafe" that was not mentioned by Nora but is important for the embedded community in particular is the use of "unsafe" to flag things which from the point of view of Rust itself are fine, but are dangerous enough to be worth having a human programmer directed away from them unless they know what they're doing. From Rust's point of view, "HellPortal::unchecked_open()" is safe, it's thread-safe, it's memory safe... but it will summon demons that could destroy mankind if the warding field isn't up, so, that might need an "unsafe" marker and we can write a checked version which verifies that the warding is up and the stand-by wizard is available to close the portal before we actually open it, the checked one will be safe.
Something that would make it harder to water down the meaning of the clearly defined unsafe keyword to suddenly mean something else.
Using the unsafe keyword to mark a function as "potentially dangerous" is just wrong.
Just prefix your functions with something like "dangerous_call", but don't misuse unsafe!
I have several times in code review prevented people from marking safe interfaces as "unsafe" because they are "special and concerning", overloading the usage of unsafe is itself dangerous.
For an example, consider Vec::set_len in the standard library. Which only contains safe code, but lets you access uninitialized memory and beyond the length of your allocation by modifying the length field of vector: https://doc.rust-lang.org/src/alloc/vec/mod.rs.html#1264
You might be able to fix this with a lint that looked at a bit more context though, `unsafe fn foo()` in a module (or even crate) with no actually unsafe operations is very likely wrong. Likewise `unsafe fn foo()` which performs no unsafe operations and only accesses fields, statics, functions, and methods that are public.
self.len = new_len;
No unsafe operations in sight. Should the compiler emit a warning here?Seems like you could create a sort of userland-unsafe by using a closed trait[1] and requiring its use on a 'dangerous' method or struct or whatever:
mod my_unsafe {
pub trait MyUnsafe {}
}
pub struct MyUnsafe;
impl my_unsafe::MyUnsafe for MyUnsafe;
pub fn dangerous_function<U: my_unsafe::MyUnsafe>() {}
fn main() {
dangerous_function::<MyUnsafe>()
}
Obviously this doesn't let you do a block of unsafe without having to repeat it like `unsafe {}` does, but it doesn't leave you much room to do the dangerous things without the shrinkwrap agreement either (and turbofish are so ugly at least for me they'd be a deterrent).That said I find the named use-case kind of weird. The whole point of the library is to do these unsafe things, so it's kind of silly to be like "don't forget it's dangerous!"
I think an interesting part of unsafe Rust is the interplay between positive and negative polarities (unsafe/safe fn vs unsafe blocks basically) and I think that adding new syntax is the way to leverage this kind of idiom in the new kind of unsafe
The safety benefits of Rust appear when you aren't willing to formally prove all the code you write correct. That's because you only have to prove the unsafe code, plus the memory model, correct in order to guarantee memory safety for the entire program. This is less burdensome than in C, where to do the same you have to prove correctness properties about the specific code that makes up the whole program. Rust makes it easier to prove certain properties about the program (far easier in the case of memory safety), but it was always possible.
In particular, the dirty secret of C verifiers is that they don’t handle pointers all that well. Either you find yourself doing a lot of manual proof work or you have to dramatically simplify the memory model.
In contrast, when verifying safe Rust, the rules of the borrow checker allow us to dramatically simplify the verification work. All of a sudden verifying a manual memory program with pointers (borrows) becomes as simple as verifying a basic imperative language. I’ve been working on a tool: https://github.com/xldenis/creusot to put this into practice
On the other hand, the moment you dive into unsafe, all bets are off and you find yourself wading through the marshes of (weak) memory models with your favorite CSL as your only friend.
Note that there are other tools trying to deal with formal statements about Rust programs. AIUI, Rust developers are working on forming a proper team or working group for pursuing these issues. We might get a RFC-standardized way of expressing formal/logical conditions about Rust code, which would be a meaningful first step towards supporting proof-carrying code directly within Rust.
Not without a formal model of C, and the C standard is only an informal, natural-language text. Rust having memory safety as its express goal (which is basically table stakes for any sort of workable language semantics) means that it's at least realistic to think about a formal semantics for Safe Rust. Then you "just" need to deal with the uncomfortable reality that lots of Safe Rust facilities actually bottom out into Unsafe Rust, which is why it turns out you must care about its semantics too. But the hope is ultimately that the small portion of real-world Rust codebases that's Unsafe Rust might not "infect" the Safe Rust to the point of making verification as practically unworkable as in C.
It seems like it is a long way off from being possible to do that for rust.
See https://doc.rust-lang.org/book/ch19-01-unsafe-rust.html
Check out also the "Rust considers it safe to" section of https://doc.rust-lang.org/nomicon/what-unsafe-does.html
For example, String::from_utf8_unchecked[1]. There's a comment in the documentation trying to justify why invalid strings could cause memory unsafety, but its pretty weak:
> This function is unsafe because it does not check that the bytes passed to it are valid UTF-8. If this constraint is violated, it may cause memory unsafety issues with future users of the String, as the rest of the standard library assumes that Strings are valid UTF-8.
Like, sure, we could make that argument for any runtime invariant. "Look, this is totally memory safe but if you call this method we assume this constraint is valid. We're marking this unsafe anyway because we want people to take care in using it. So uh, its memory-unsafe because violating this constraint could cause memory unsafety issues in the future or something? Yeah that'll do."
It seems like lots of folks quietly want unsafe to mean "this method skips runtime checks", with no specific grounding in memory-unsafety. And the line between those two ideas is super blurry in practice, even in the standard library.
[1] https://doc.rust-lang.org/std/string/struct.String.html#meth...
[1] https://github.com/rust-lang/rust/blob/027a232755fa9728e9699...
I see enforcing utf8 as about on par with enforcing that a boolean be 0 or 1, and not 5.
How do you feel about this list of what's unsafe and undefined? https://doc.rust-lang.org/nomicon/what-unsafe-does.html
UTF8 correctness is an invariant of the code. Its not something the compiler understands or cares about it.
But really, I get it! My point is that rust is rife with people marking functions "unsafe" when they mean a method will cause bugs / "runtime UB" if used improperly. Even the standard library does this, so it feels a bit rich to cry foul when 3rd party libraries (ab)use "unsafe" in the same way.
I certainly use unsafe like that, in order to mark methods which, if misused, will violate internal invariants of my code.
How does that list of unsafe features justify the unsafe marker on String::from_utf8_unchecked?
If I had a way to set some booleans "unchecked", in a way that could make the numerical value be 5, would you label it unsafe?
I really think it should be unsafe, because it's crazy to make ifs or switches on booleans be non-exhaustive, and it's also crazy to have an "invalid boolean" path anywhere one is used.
When you have an encapsulated value, with special accessors to make sure it's always in range, then you don't have to have an unchecked setter function at all. But if you do add one, I think it makes sense to make it unsafe. That function is basically allowing a blind memcpy over a piece of data. It ties into safety pretty strongly. That kind of access is a lot like dereferencing a raw pointer, and not making it unsafe means you allow a lot of those "invalid value" cases in the link to be possible in safe code.
The Rust compiler is allowed to (and sometimes required to) use the invalid values of a child value to represent other cases of an enum. For example, this is how it optimizes `Option<&T>` into "pointer where null is None and any other value is Some".
This would technically also allow it (I don't know whether it currently does this..) to optimize the following enum:
enum Foo {
A(bool),
B(bool),
C(bool),
}
into the following single-byte representation: A(false) => 0,
A(true) => 1,
B(false) => 2,
B(true) => 3,
C(false) => 4,
C(true) => 5,
With this representation, `A(5 as bool)` would be impossible to distinguish from `C(true)`, which would be complete and utter nonsense.Yes, that's why I used it as the example.
And you don't need anything that complicated, either. A simple if or switch statement might be optimized by the compiler for 0 and 1 and jump into random code for 5.
I'd be happy to extend the definition of "unsafe" to mean "If you misuse this, you may violate some internal invariants of the library. And that may cause unexpected bugs."
Thats a broader, but in my opinion much more clear definition of unsafe, which covers how the keyword is actually used in actual code. (In std::String and elsewhere).
I've implemented high performance b-tree and skiplist implementations in rust. There's plenty of functions in both libraries which I've marked as unsafe because if you use them carelessly, you'll violate some internal invariants. Will the library break as a result? Yes. Will the resulting bugs include memory corruption? I really have no idea, and I don't really care enough to go in and test that. So I've marked the methods unsafe, and provided safe APIs for consumers to use instead.
Do you think this is an appropriate use for unsafe? If not, I'm curious to hear why.
Or more simply: Worry about all kinds of undefined behavior. That makes many issues easier to find, and you don't need to chain on additional logic to figure out how it might corrupt memory.
One difference is that Rust has a separate type for String-like objects that do not respect the invariant, namely Vec<u8>. So those who wish to write memory-safe code acting on arbitrary bytes can simply use that. Also, it's quite normal for an invariant defined entirely in Safe Rust to impact memory safety in a very real sense; consider Send and Sync. These are seemingly arbitrary labels, but their semantics is nonetheless quite well defined.
Rust gave up this notion when they decided that mem::forget could, in fact, be safe, since even though it was initially marked unsafe it could be trivially implemented in safe code using only safe constructs from the stdlib.
We need another keyword or concept to refer to unsafeness that isn't related to UB. Some people suggested tagging unsafe with something like
// declare SummonsDemons as some kind of safety market that goes beyond preventing UB
unsafe(SummonsDemons) fn f() {
// summons demons
}
fn safecode() {
// SAFETY: the demons are cool this time
unsafe(SummonsDemons) { f() }
}I thought there were subtle language differences such that if you took ordinary, perfectly valid as safe rust code and marked it as unsafe the result could be incorrect? Am I mistaken-- maybe that was some pre-1.0 property of the language and it's actually okay to go peppering around unsafe for typechecking like usage?
Separately, if people are creating unsafe interfaces to protect them against uses that fail to uphold their required invariants it would probably be really useful if they could be tagged with specific required unsafty-capabilities such that if a function requires foo-stability and bar-stability you don't accidentally call it with code that only guarantees foo-stability under a mistaken impression that foo-stability was all the unsafty of the function was related to.
For example volatile_register is a crate for representing some sort of MMIO hardware registers. It will do the actual MMIO for you, just tell it where your registers are in "memory" and say whether they're read-write, read-only, or write-only just once, and it provides the nice Rust interface to the registers.
https://docs.rs/volatile-register/0.2.1/volatile_register/st...
The low-level stuff it's doing is inherently unsafe, but it is wrapping that. So when you call register.read() that's safe, and it will... read the register. However even though it's a wrapper it chooses to label the register.write() call as unsafe, reminding you that this is a hardware register and that's on you.
In many cases you'd add a further wrapper, e.g. maybe there's a register for controlling clock frequency of another part, you know the part malfunctions below 5kHz and is not warrantied above 60kHz, so, your wrapper can take a value, check it's between 5 and 60 inclusive and then do the arithmetic and set the frequency register using that unsafe register.write() function. You would probably decide that your wrapper is now actually safe.
So, this isn't a case that the `unsafe` is there just as a warning lint, it has an actual meaning, protecting the memory safety invariants.
That was what I meant, thanks for the answer! Though if you have an example of the other thing I'd be open to that too
"Rip and tear, until it's done!"
"The only thing they fear is you"
playing in the background.
Like this: https://www.infoq.com/news/2021/11/rudra-rust-safety/, and I quote: "In C/C++, getting a confirmation from the maintainers whether certain behavior is a bug or an intended behavior is necessary in bug reporting, because there are no clear distinctions between an API misuse and a bug in the API itself. In contrast, Rust's safety rule provides an objective standard to determine whose fault a bug is."
People bring up `unsafe` Rust as an argument against the language, but to me it appears to be an argument `for` it.
Those people usually don't understand what `unsafe` is.
> The biggest failure in Rust‘s communication strategy has been the inability to explain to non-experts that unsafe abstractions are the point, not a sign of failure.
Appeared on TWiR: https://this-week-in-rust.org/blog/2021/10/20/this-week-in-r....
"I started learning Rust in 2015. I've been using it since 2016. I've been a compiler team member since 2017. I've been paid to write Rust code since 2018. I have needed to use `unsafe` <5 times. That's what why Rust's safety guarantees matter despite the existence of `unsafe`."
For me, I was being a bit lazy with loading some config on program startup that would never change, so I used a `static mut` which requires `unsafe` to access. Turns out I was able to figure out a way to pass my data around with an `Arc<T>`. I think either way would have worked, but I figured I should avoid the unsafe approach anyway.
https://github.com/shepmaster/stack-overflow-relay/blob/273a...
[1] https://stackoverflow.com/questions/10529284/is-there-ever-a...
hold_my_beer {
}}
My problem with the word “unsafe” is that many people seem to think unsafe code blocks necessarily have security issues. Hence the bullying of library authors even in cases where the authors did their homework.
Anything about the mechanics of the implementation is a second-order detail compared to that, as far as I’m concerned.
Also, I dislike the word "unsafe" since "unsafe code" is easily (mis)interpreted to mean "invalid/UB code", but "invalid/UB code" is officially called "unsound" rather than "unsafe". Unsafe blocks are used to call unsafe functions etc. (they pass information into unchecked operations via arguments), and unsafe trait impls are used by generic/etc. unsafe code like threading primitives (unsafe trait impls pass information into unchecked operations via return values or side effects).
unchecked would make a good keyword name for blocks calling unchecked operations, but I'm not so sure about using it for functions or traits.
For example, this is UB:
unsafe { std::hint::unreachable_unchecked() };
And this is an unsound code, that may not contain any UB if I call `unsound()` with `false`: fn unsound(trigger: bool) {
if trigger { unsafe { std::hint::unreachable_unchecked() }; }
}The real issue here is Rust fans trying to make Rust into a type dependent language while abusing unsafe's original purpose.
It would be pretty interesting to me if someone wrote a survey of what different languages consider to be "unsafe", including specific operations like this. For example, it looks like "sizeof" is relegated to the "unsafe" package in Go, which strikes me as strange.[1] I'd love to read a big comparison.
Some recent developments: https://gankra.github.io/blah/tower-of-weakenings/
Is it your expert opinion that Wirth foresaw the problems we have writing this bit-banging code on modern hardware and solved them correctly despite never formalising what they actually are ?
Lists of languages that don't even try to do what Rust is doing here suggest you haven't thought very hard about even what the problem is. Of course if you aren't sure what the problem is it may be trivial to convince yourself that you've "solved" it.
Rust is following Ada's footsteps for a newer generation, while still not having something comparable to SPARK.
So what it is Rust trying here?
C# is subject to the whims of the host CLR for both type safety and memory model, and as a result is inherently unsafe on CLRs that actually exist (primarily, Microsoft's) because dangerous was much faster.
They're both dramatically better than C++ and so give us an opportunity to assess Herb's claim that if C++ was 90% safer nobody would bother him about how ludicrously dangerous it is. It does seem like people are comfortable in C# and Go as a result. But, once in a while something very strange happens and it gives these developers no comfort to learn that the Undefined Behaviour was foreseen and they get a WONTFIX back from the language's implementers (e.g. C# can't conceive of how two booleans can both be true yet not equal, however the CLR does not guarantee what C# wants here so C# loses).
So that's what Rust is for. Mostly though, Rust is just very nice to use. I don't actually need Aria's Tower of Weakenings, most of the time I don't need even need unsafe, but I'd still rather be in Rust.
As for data races, great that it fixes data races among threads, yet it does nothing against shared resources in distributed computing, which are what really matter in the age of microservices and protection against spectre like attacks.
What is "it" here which you believe has been advised against? Please try to be specific.
> As for data races, great that it fixes data races among threads
Yes, it is pretty good. It was pretty great the last few times you "forgot" about this too.
"In addition, unsafe does not mean the code inside the block is necessarily dangerous or that it will definitely have memory safety problems: the intent is that as the programmer, you’ll ensure the code inside an unsafe block will access memory in a valid way."
Although it appears I stand corrected in regards to Rustonomicon, as it seems to adopt the position that unsafe should apply to more than memory corruption bugs and failed to find the sentence I was looking for.
> Yes, it is pretty good. It was pretty great the last few times you "forgot" about this too.
I never forget about it, from my point of view they are oversold, and will keep arguing they aren't enough for distributed computing deployment scenarios with heavy use of acess to process external shared resources.
Fearless concurrecy isn't something I care about when accessing database rows from multiple threads across database connections, writing to NUMA regions across process and so forth, shared storage,....
Races are still bound to happen in such scenarios, if care is not taken, and fearless concurrency is of little help in such cases.
> - Dereference a pointer that does not point to the type the compiler thinks it points to. [...]
> - Cause there to be either multiple mutable references or both mutable and immutable references to the same data at the same time. [...]
Sometimes, I wish there was a systems language that let you exactly do that. Dereference any memory location and see what memory is there. Use a cast to reinterpret the bits. If the memory address is unmapped, then maybe just throw an exception. Allow multiple threads to fill in different spots in a huge array at the same time. Basically write C as high level assembler, like you did in the 90s.
I know -fno-strict-aliasing and some other flags can get you mostly there, but you are still violating the spec and it is at best tolerated. I also know that you loose some optimization potential, but it also reduces UB and that is a trade off you should be able to choose.
“Safety, in Rust, is very well-defined; we think about it a lot.”
You can’t just construct some elaborate static analysis complex and call your language “safe” as a result, that’s a surefire way to make your language impenetrable to anyone outside of your field of expertise.
If Rust didn't have unsafe, it could have that feature. Instead, the Rust compiler assumes a lot of data is !Sync, such as anything that might indirectly contain a trait object (unless explicitly given + Sync) which might contain a Cell or RefCell, both of which are only possible because of unsafe.
Without those, without unsafe, shared borrow references would be true immutable borrow references, and Sync would go away, and we could have Seamless Concurrency in Rust.
I often wonder what else would emerge in an unsafe-less Rust!
Still, Given Rust's priorities (low level development) and the borrow checker's occasional need for workarounds, and the sheer usefulness of shared mutability, it was a wise decision for Rust to include unsafe.
[0] https://verdagon.dev/blog/seamless-fearless-structured-concu...
Rust devs have thought about implementing "true immutable" before and found it to be problematic. It would come in quite handy for mostly anything related to FP or referential transparency/purity, but these things turn out to be very hard to reconcile with the "systems" orientation of Rust. Perhaps the answer will reside in some expanded notion of "equality" of values and objects, which might allow for trivial variation while verifying that the code you write respects the same notion of "equality".
And even without unsafe, you still couldn't assume all data is Sync. Counter-examples include references to data in thread-local storage, and most data used with ffi.
CPython is also less aggressive about adopting mitigation strategies compared to Rust, although it has improved a lot in the last few years.