Undefined vs. Unsafe in Rust
manishearth.github.io
manishearth.github.io
The `unsafeXYZ` idiom is very common in the functional programming world. For individual functions, prefixing them with `unsafe` generally means, this function could result in undefined behaviour. However, there is the other idea that we mark something as `unsafe` to ask the compiler not to check it and just to trust us (`unsafePerformIO` is often used this way). Maybe having a second term would make this clearer.
However, that superset contains things that are not checked. This has no effect on any other code though; safe code with a superfluous unsafe block acts 100% exactly the same.
Does Rust have a way to prevent this class of errors? What ensures that all units have the same view of a type?
At a lower level, symbols are mangled to include a crate identification hash. This is more of a way to allow multiple versions to coexist though, as the problem is already solved by that point.
1. Compile a library libA that references type T in CrateX.
2. Add a field to T, and recompile CrateX.
3. Compile an executable a.out that passes a T to libA.
Nobody bumped CrateX's version or recompiled libA. Won't this crash? Or does this produce a "different version" of CrateX and so it will refuse to link?
pub struct S {
a: usize,
}
pub fn foo(s: &S) -> usize {
s.a
}
…as a dylib, then added another field to S and recompiled. I expected it to change the symbol names (as shown by nm), but it didn't. The metadata might be different (I don't know how to display it), but the dynamic linker doesn't know about that.However, Cargo will still ensure different versions of the same crate have different symbols, by passing its own hashes to rustc via -C metadata.
The metadata just contains info on all the types, and their hashes (or something like that), so if stuff doesn't match you'll know.
This is generally visible from the "expected type Foo but found type Foo" error, which will often mention you have two versions of the same crate.
Worth mentioning that unlike C++ or C Rust doesn't have a global name mangling scheme; so it is totally ok for two crates to have a toplevel struct Foo (unlike in C++ where you are forced to namespace them with uniquely-named namespaces). This has the side effect of it being totally ok to link two versions of the same crate together; and Rust will just complain if you try to mix the types.
Distinguishing between crates is done solely through their name and -C metadata values provided by Cargo.
Once crates are loaded by the compiler based on either their explicit path (via --extern) or by being a dependency of another dependency (and there the name and -C metadata prevent collisions), "two versions of the same crate" appears no different than "two different crates".
The Rust compiler tends to "index" information (e.g. turning strings into various IDs) as quickly as possible, so a lot more semantics are "by identity" than "by syntax", and that helps when multiple identities may share a name.
That includes compiling against already compiled crates, instead of header files you have serialized semantic types and functions, which all use proper identities to "name" anything they use in turn - you never have to be looking for a definition, or risk using the wrong one.
I used these sources:
// a.rs
#![crate_type = "lib"]
pub struct Foo { pub x: i32 }
// b.rs
#![crate_type = "lib"]
extern crate a;
pub fn foo() -> a::Foo { a::Foo {} }
// c.rs
extern crate b;
fn main() {
println!("{}", b::foo().x);
}
The field `x` from crate `a` was originally not there, and added after compiling crate `b`. The error when trying to compile crate `c` was: error[E0460]: found possibly newer version of crate `a` which `b` depends on
--> c.rs:1:1
|
1 | extern crate b;
| ^^^^^^^^^^^^^^^
|
= note: perhaps that crate needs to be recompiled?
= note: crate `a` path #1: liba.rlib
= note: crate `b` path #1: libb.rlibIn practice you compile everything with Cargo and it ensures all affected crates are recompiled when interfaces or settings change, so everything matches.
i guess the reason behind using the c-structs is to avoid manually calculating the offsets for N different ruby versions you want to support. i guess the alternative to casting/serialization would be to extract the offsets from the generated rust bindings which is apparently unsafe in rust as well. heh
Depending on your sitaution, you don't have to use unsafe yourself; for example, the byteorder crate can help here, or any of the serialization libraries that can serialize to a binary format.
Code in an unsafe block may generate undefined behavior if it is used incorrectly, but, if used correctly, will not.
Rust code which is not in an unsafe block cannot generate undefined behavior (aside from bugs in the compiler/underlying libraries).
> Code in an unsafe block may generate undefined behavior if it is used incorrectly, but, if used correctly, will not.
The point of this was that code in an unsafe block should not be able to generate undefined behavior no matter how it is used from safe code, otherwise that unsafe block is unsafe, not safe.
> Rust code which is not in an unsafe block cannot generate undefined behavior
That's false, code outside of an unsafe block can generate undefined behavior by calling unsafe code that was not written safely.
> tl;dr
Please don't post misleading summaries on nuanced topics.. especially when the original post is short.
Is there any way to write safe code in Rust then?
To guarantee complete rust-safety, modulo LLVM bugs, do not use any Rust code that has unsafe blocks in it (this includes any code that has memory allocations, such as Vec).
Less strictly one can refuse to use code with unsafe blocks that has not been thoroughly vetted by experts. Things like Vec may therefore be used depending on how strict one's definition of "vetted by experts" is, but rolling one's own "unsafe" block code is not allowed.
I write rust-safe code all the time using the second standard, even though under the covers it's littered with unsafe.
So, while it's possible to manually check the use of unsafe, there is no way I can currently do it automatically, right? I noticed there is a #![forbid(unsafe_code)], which I can add to my source file, but this does not check that whatever functions I am importing or calling do not underlyingly use unsafe, right?
Moreover, I admit I am being lazy here, but any idea whether there are sub-communities in Rust that are keen on creating these kinds of safe programs?
People have varying tolerances on the use of unsafe code. I'm not aware of any coherent sub-community with a fundamentalist take on it though.
My own personal tolerance is "when using it, carefully justify it." Generally, this takes the form of making the code run faster. Otherwise, I'll sacrifice almost everything else to avoid unsafe. Thankfully, said sacrifices are neither common nor arduous (in my experience).
I sometimes wonder if we should maybe invest more in non-turing-complete language research. Obviously Haskell (and Rust) are improvements to the status-quo (critical code can be guarded explicitely via unsafe attributes and written by experts), but maybe this does not go far enough.
If the 'unsafe' blocks contain their unsafety correctly (thus being 'safe' 'unsafe' blocks using the two different meanings of safe the article discusses), then using them is fine. The rust stdlib has a promise that all of their 'unsafe' blocks not marked as unsafe are 'safe', so using the stdlib you're writing 'safe' code (though it's possible it's unsafe if there are stdlib/compiler bugs).
As the article explains, it's very nounced and, while imperfect, still a damn sight better than C++.
https://github.com/rust-lang/rust/issues/33813
This basically exposes LLVM's poison semantics (a kind of deferred UB) to Rust code:
https://llvm.org/docs/LangRef.html#poisonvalues
Code in another module that doesn't even know it is calling an unsafe function could trigger UB by using this.
It seems like you'd still have to dereference it, which is unsafe, yes?
I haven't tried very hard to write a seemingly safe Rust program using ptr:offset that LLVM will optimize into a crash. It's probably possible with a bit of work; this is basically C's signed overflow UB in a different guise, and programmer misunderstanding of that has caused serious bugs and security issues over the years.
Edit: There are some other known ways of causing UB from safe Rust code, but they’re considered bugs - and some longstanding ones are on the way to being fixed in the next few months, thanks to MIR borrow checking and saturating float->int casts.
#[link_section = ".data"] fn main() { }(link_section might be counted as one of these. or just tweaking your link flags)