Cve-rs: Fast memory vulnerabilities, written in safe Rust
github.com
github.com
(Have you ever looked at a fuzzer-generated failing piece of data and not though "oh whoa good point". That's interesting!)
- https://github.com/pkolaczk/latte/blob/main/src/main.rs
- https://github.com/pkolaczk/latte/blob/main/src/exec.rs
There was one fundamental "aha" moment for me when it clicked: move semantics. Once I learned it, suddenly 99% of stuff became simple (including making async quite nice really, contrary to popular HN beliefs).
I would also generally agree that a single pass through the official book should mostly be enough to be able to read and understand the op.
I’ve written at least two other non-toy Rust programs and I haven’t hit “lifetimes everywhere” problem there either. I guess the most lifetime related problems stem from trying to program the Java/Python/JS style in Rust - from overusing references and indirection. If you model your program as a graph of objects referencing each other then I can imagine you have a hard time in Rust.
[,{a,'z':q|0}[`${a?.x??b}`]]I mean, a read through the guide makes this very readable, but it doesn’t change that it looks like symbol soup and have lots of awkward repetitive strain to type it out.
I have a custom layout on my keyboards which makes each symbol reachable without moving my hands, which made it a lot more pleasant.
`'a` is a label for a memory area, like goto labels in C but for data. You can read it as (memory) pool A. Roughly, when function entered, memory pool is created, then destroyed at exit.
`'static` is special label for static data (embedded constants).
`()` is nothing, like void in C.
`&()` is an reference to nothing (an address) like `void const *` in C.
`&&()` is an reference to reference to nothing, like `void const * const *` in C.
`& 'static & 'static ()` is like void `const * const *` to a built-in data in C.
`& 'a & 'b ()` tells compiler that second reference is stored in pool 'a, while data is stored in pool 'b. (First reference is in scope of current function.)
`_` is a special prefix for variables to instruct compiler to ignore warning about unused variable. Just `_` is a valid variable name too.
static UNIT: &'static &'static () = &&();
fn foo<'a, 'b, T>(_: &'a &'b (), v: &'b T) -> &'a T { v }
Let's say that `&'a` is from a function `a()`, while `&'b` is from a function `b()`, which called from the function `a()`.The trick here is that we relabel reference from pool `'b`, from an inner function, to pool `'a` from outer function, so when program will exit from function `b()`, compiler will destroy memory pool `'b`, but will keep reference to data inside until end of the function `a()`.
This should not be allowed.
I guess stated another way, I don't generally have issues reading code from a wide swath of languages. If someone plopped me in front of a rust codebase I'd be at the mercy of the manual for quite a long time.
Thank you again, sincerely.
& is used for references in a lot of languages.
() is a tuple, same syntax as Python
'a isn't even novel: it was taken from OCaml, which uses it for generic types. Lifetimes are generic types in Rust, so even though it's not for an identical thing, it's related enough to be similar.
_ to ignore a name for something has a long tradition in programming, sometimes purely as a style thing (like _identifier or __identifier), but sometimes supported by the language.
fn name(args) -> {} as function syntax is not SUPER unusual. The overall shape is normal, though the choice between "fn," "fun," "func," or "function" varies between languages. The -> comes from languages like Haskell.
<> for generics is hotly debated, but certainly common among a variety of langauges.
the "static name: type = value;" (with let instead of static too) syntax is becoming increasingly normalized, thanks to how it plays with type inference, but has a long history before Rust.
So this leads to a very interesting thing, where like, it's not so much that Rust's syntax is entirely alien in the context of programming language syntax, but can feel that way unless you've used a lot of things in various places. And that also doesn't necessarily mean that just because it's influenced from many places that this means it is coherent. I think it does pretty good, but also, I'd refine some things if I were making a new language.
Which parts would you change?
I do think there's some good criticism of doing this, though, and so even a year later it's not clear to me it's a pure win.
I am sympathetic to the vague calls to action by Aria et al. to improve the syntax for various unsafe features. I understand why we ended up where we ended up but I think overall it was a mistake. I am not sure I agree with her specific proposals but I do agree with the general thrust of "this was the wrong place to provide syntactic salt." (my words, not hers, to be clear)
Ideally `as` wouldn't be a thing. Same deal, this is just one of those things that's the way it is due to history, but if we're ignoring all that, would be better to just not have it.
I am still undecided in the : vs = debate for structs.
I am sad that anonymous lifetimes in structs died six years ago, I think that would be a massive help.
Probably tons of other small things. Given the context and history, I think the Rust Project did a great job. All of this is very minor.
The = on functions idea is particularly interesting by comparison with Scala, which has shifted towards requiring = over time.
Therefore, it ends up being a really huge thing, with lots of design space. This is combined with the fact that
> they would also provide an easy, unambiguous way to "overload" functions
Not everyone sees this as a good thing.
Being slightly controversial, plus being really large, plus there being a lot of other things to do, would make me surprised if they ever land.
Here's a link from eight years ago with a link to lots of other related proposals: https://internals.rust-lang.org/t/pre-rfc-named-arguments/38...
In particular, explicit use of lifetimes tends to be seen as somewhat more of a last resort, although the language requires it more frequently than I like. Furthermore, the soupiest of the syntax requires the use of multiple layers of indirection for references, which itself tends to be a bit of a code smell (just like how in C/C++, T * tends to be somewhat rare).
The real question what is the use of Rust for you. Do you work on anything where Rust could be a value?
This is not a representative sample of Rust. That's explicitly triggering edge cases which requires abuse of syntax you wouldn't normally see.
Check out this for something more realistic that anyone should understand https://github.com/ratatui-org/ratatui/blob/main/examples/ca...
And even if not, this example isn’t particularly hard to decompose and understand once you have a basic grasp of the independent underlying bits (generics, lifetimes, borrows). It’s just combining all of them pathologically into an extremely terse minimal reproduction.
It’s like saying you could never understand Java because someone linked to an AbstractClientProxyFactoryFactoryFactoryBean.
Just because "int ((foo)(void ))[3]" is a valid C declaration doesn't mean that all of C code looks like that.
C++ has the same level of fuckery without memory safety. I don't think there is an extreme level of fuckery in Rust. I wish they got more influence from the ML languages than C++ but it is not unbearable.
But I guess there is another trick.
Feels a bit like early chess engines. They tried to create super sophisticated heuristics to determine move quality, but then one person (or more probably multiple people independently more or less simultaneously) realized it's both easier and better to just do more or less the simplest thing that can possibly work, but do it as much as possible in the allotted time.
In other words - don't try to logically decide on the best move with some super advanced set of rules. Just assign a fairly basic and straightforward score based on calculating every reasonble move out 15, 20, 30 moves deep.
(I suppose chess engines also have an end goal of perfect play, but that is a much harder problem)
Saying "just do the thing correctly, duh!" is easy. Ya know the saying about the difference between theory and practice?
Both type A engines (to use Shannon's terminology) and type B engines are relatively weak. The strong engines are all hybrids.
Just because the compiler as a whole isn't the fastest one around doesn't mean the responsibility falls equally on every constituant part of the compiler.
There's a new borrow checker however, but that's not going to fix this either
“Similarly, the reason why niko's approach is not yet implemented is simply that getting the type system to a point where we even can implement this is Hard. Fixing this bug is blocked on replacing the existing trait solver(s): #107374. A clean fix for this issue (and the ability to even consider using proof objects), will then also be blocked blocked on coinductive trait goals and explicit well formed bounds. We should then be able to add implications to the trait solver.
So I guess the status is that we're slowly getting there but it is very hard, especially as we have to be incredibly careful to not break backwards compatibility.”
https://blog.rust-lang.org/inside-rust/2023/07/17/trait-syst...
“The new trait solver implementation should also unblock many future changes, most notably around implied bounds and coinduction. For example, it will allow us to remove many of the current restrictions on GATs and to fix many long-standing unsound issues, like #25860. Some unsound issues will already be fixed at the point of stabilization while others will require additional work afterwards.”
Apparently there's 84 open issues which the Rust developers consider unsound issues. https://github.com/rust-lang/rust/issues?q=is%3Aissue+is%3Ao...
That's it. Rust's claims are over and done with. How can it be safe if it's unsound? By the principle of noncontradiction and the laws of thought itself, let no one speak that word anymore.
You can call Rust memory safer, but until those bugs get fixed, it's wrong to call it safe.
Just like that not every security vulnerability is equally fatal, not every soundness bug is equally fatal. I reckon about three levels of severity: inherent to the design itself, not inherent to the design but reasonably user-visible, and pathological. As pcwalton pointed out, miri does show that this particular soundness bug is NOT inherent to the language design, and I believe that's true for most of 84 unsound bugs (please let me know any counter-example though, I haven't fully checked them). It remains to be seen whether there exist soundness bugs that are still user-visible enough.
> Rust’s rich type system and ownership model guarantee memory-safety and thread-safety — enabling you to eliminate many classes of bugs at compile-time.
Note that it explicitly mentions "memory" safety and "thread" safety, which has a specific but reasonable definition in Rust (for example, memory safety doesn't cover physical memory leak). Also it explicitly mentions which portion of Rust is responsible for such guarantees, namely "type system" and "ownership model", and it is reasonably claimed that the ideal implementation of both will indeed completely achieve such guarantees. To be clear, the current implementation is also very close to that ideal to make such claim meaningful, but there is always a difference between the ideal and the practice.
Rust built its reputation around the idea that they can crush security bugs by making them impossible. They should be holding themselves to a higher standard than that "in practice" leeway. If a malicious actor can tease Rust into behaving in a way that contradicts its safety guarantees, then it could be serious.
Maybe your corporate policy is to configure Rust to allow zero unsafe code. Some crate you're depending on gets hijacked. It uses the cve-rs to crash your system even though Rust says it's 100% safe code.
The safety in a programming language is mostly protecting the programmer against itself. The probability for a programmer to write this kind of code by mistake is close to zero, as opposed to UB in C or C++ that are pretty common. To make a vulnerable program with this kind of issue, the programmer would have to make them on purpose, what is unlikely unless for this kind of joke repository.
Yes, because they were caught in review and tests, or were patched in a bugfix release before being widely exploited. Rust catches safety issues during compilation, before you test or commit.
> Some crate you're depending on gets hijacked
Rust's safety is meant to protect against a certain class of programmer mistakes. You still need to audit your dependencies and sandbox untrusted code; the language designers have never claimed otherwise.
A large portion of the bugs in that list require unstable language features. Most of the remainder are codegen bugs (miscompilations, ABI mismatches, linker problems, etc). The list of core type system soundness bugs is a lot shorter, it's tracked here: https://github.com/orgs/rust-lang/projects/44
That bug is marked as I-unsound, which means that it introduces a hole in the type system.
And so are all other similar bugs, i.e., your concern seems to be unfounded, since you can actually click on the I-unsound label, and view all current bugs of this kind (and past closed ones as well!).
To be clear I wasn't trying to imply the rustc maintainers were ignorant of the difference. I meant that Rust programmers seem to treat fundamental design flaws in the language as if they are temporary bugs in the compiler. (e.g. the comment I was responding to) There's a big difference between "this buggy program should not have compiled but somehow rustc missed it" and "this buggy program will compile because the Rust language has a design flaw."
I am not at all familiar with Miri. Does Miri consider a slightly different dialect of Rust where implicit constraints like this become explicit but inferred? Sort of like this proposal from the GH issue: https://github.com/rust-lang/rust/issues/25860#issuecomment-... but the "where" clause is inferred at compile time. If so I wouldn't call that a "fix" so much as a partial mitigation, useful for static analysis but not actually a solution to the problem in rustc. I believe that inference problem is undecidable in general and that rustc would need to do something else.
fn foo<'a, 'b, T>(_: &'a &'b (), v: &'b T) -> &'a T { v }
should be equivalent to fn foo<'a, 'b, T>(_: &'a &'b (), v: &'b T) -> &'a T where 'b: 'a { v }
because the type &'a &'b is only well-formed if 'b: 'a. However, in the implementation, only the first form where the constraint is left implicit is subject to the bug: the implicit constraint is incorrectly lost when 'static is substituted for 'b. This is clearly an implementation bug, not a language bug (insofar as there is a distinction at all—ideally there would be a written formal specification that we could point to, but I don’t think there’s any disagreement in principle about what it should say about this issue).The point is that rustc does not even implicitly have the "where 'b: 'a { v }" clause. The programmer knows "where 'b: 'a { v }" is true because otherwise &'a &'b would be self-contradictory, and in most cases that's more than enough. But in certain edge cases this runs into problems with the contravariance requirement since we can substitute either one of the 'bs with a type that outlives 'b (i.e. 'static), and the design of contravariance lets us do this pretty arbitrarily.
Look closely at what happens:
fn foo<'a, 'b, T> (_: &'a &'b (), v: &'b T) -> &'a T
There is an implicit "where 'b : 'a" clause at the end of this declaration - this clause would be explicit if Rust was more like OCaml. The reason this clause is there implicitly is that in correct Rust code you can't get &'a &'b if 'b : 'a doesn't hold. So when a human programmer reasons about this code they implicitly assume 'b: 'a even though there's nothing telling rustc that this must be the case.This runs into a problem with the requirements of contravariance, which let us replace any instance of 'b in foo with any type 'd such that 'd : 'b, in particular 'static:
fn foo<'a, 'b, T> (_: &'a &'static (), v: &'b T) -> &'a T
Since contravariance allows us to replace the &'a &'b with &'a &'static without changing the second &'b, the completely implicit constraint 'b : 'a is no longer actually being enforced, and there's no way for it to be enforced. This is not a compiler bug! It's a flaw in the design.In particular I don't think there's anything especially magical about 'static here except that it works for any type 'a[1]. If you had a specific type 'd such that 'd : 'a then I think you could trigger a similar bug, converting references to 'b into references to 'd.
[1] Actually maybe "static UNIT: &'static &'static () = &&();" is more critical here since I don't think &'d &'d will work. Perhaps there's a more convoluted way to trigger it. This stuff hurts my head :)
As far as I can tell, the main reason to argue "language design" vs. "compiler bug" is to imply that language design issues threaten to bring down the entire foundation of the language. It doesn't work like that. Rust's type system isn't like a mathematical proof, where either the theorem is true or it's false and one mistake in the proof could invalidate the whole thing. It's a tool to help programmers avoid memory safety issues, and it's proven to be extremely good at that in practice. If there are problems with Rust's type system, they're identified and patched.
[1]: https://github.com/rust-lang/rust/issues/25860#issuecomment-...
It's a little like making a type system turing-complete then saying, "we can fix the halting problem in a patch".
If you change the design, then you can fix it. This is what I interpret "design flaw" to mean.
Interestingly enough, in practice even mathematical proofs aren't like that either: flaws are routinely found when papers are submitted but most of the time the proof as a whole can be fixed.
Wiles first submission for his proof of Fermat's last theorem in 1993 is the best known example, but it's in fact pretty frequent.
I agree that this isn’t a “fundamental design flaw in the language”, but Miri is irrelevant to proving that.
Or to put it another way, the reason that Rust’s implied bounds issue is not a fundamental language issue is that it almost certainly can be fixed without massive backwards compatibility breakage or language alterations, whereas making C safe would require such breakage and alterations. But Miri tells us nothing about that.
I suppose the end result is the same, but it might impact any justification around whether the fix should be a minor security patch or a major version bump and breaking update.
In the case of user code that isn't unsound but breaks with the changes to the compiler/language, that would be breaking backwards compatibility, in which case there might be a need to relegate the change to a new edition.
- Assume our language has a specification, even if it's entirely in your head
- a "correct" program is a program that is 100% conformant to the specification
- an "incorrect" program is a program which violates the specification in some way
Let's say we have a compiler that compiles correct programs with 100% accuracy, but occasionally compiles incorrect programs instead of erroring out. If the language specification is fine but the compiler implementation has a bug, then fixing the compiler does not affect the compilation behavior of correct programs. (Unless of course you introduce a new implementation-level bug.) But if the language specification has a bug, then this does not hold: the specification has to change to fix the bug, and it is likely that at least some formerly correct programs would no longer obey this new specification.
So this is true:
> It's still something you fix by changing what programs the compiler accepts or rejects
But in one case you are only changing what incorrect programs the compiler accepts, and in the other you are also changing the correct programs the compiler accepts. It's much more serious.
It does not, and at the current pace, it might never have a spec.
The reason there is no spec - not even an hypothetical spec in my head - is that the exact semantics of Rust has not been settled.
With the constraints the Rust project are operating with, the only way forward I can think of is following the ideas laid out in post
https://faultlore.com/blah/tower-of-weakenings/
With the understanding that you can have multiple specs if one is entirely more permissive than the other (and as such, programmers must conform to the least permissive spec, that is, the spec that allows the smallest number of things)
But the problem is, Rust doesn't even have this least permissive spec. Or any other.
If it's been known since 2015 and not fixed, that's pretty suggestive.
That means nothing. In Rust, "everyone is a volunteer" and you're not allowed to expect things to be fixed unless you do it yourself - so the fact that this hasn't been fixed is simply an artifact of the culture, not necessarily the difficulty of the problem.
On the one hand that's encouraging - you really need to want to trigger it, it doesn't look like a bug that's likely to get accidentally written. Even then, it required a bunch of anti-optimization barriers and stuff to actually be exploitable.
On the other hand, one compiler bug and safe, well, isn't
I still like Rust.
Does it crash safely as well? I did not test it, I more than 640 KB of ram.
- [0] https://github.com/Speykious/cve-rs/blob/d51f52dd64f148a086e...
I should point out that the reason the bug is difficult seems to mainly be because of backwards compatibility concerns. If compatibility weren't an issue, Rust could just ban function contravariance and nobody would care.
* Wiping your disk
* Downloading and running code from the internet
* ...
Rust's safety guarantees don't exist to protect you from a malicious Rust programmer. They exist to protect you from mistakes a well meaning but fallible Rust programmer can make.
This greatly ignores human behavior. A LOT of people who are reviewing rust code are going to only concentrate on unsafe blocks (that's if you're lucky), because rust is advertised as "rust gives you safety", and the word "safety" strongly suggests "don't worry about these things, they are safe". if the advertised safety is a false sense of safety that's a problem.
It's also only needed for when you dip into unsafe, which is not particularly often, and so imposing this on all code would be the wrong tradeoff.
Regardless, it is not feasible for you to run your code under miri all the time, so you have to pick and choose where you use it, as a practical matter.
A thing of beauty^_^ I've seen a lot of licenses that require attribution[0], so a license that explicitly demands that you not credit the author is a lovely twist:)
[0] Read: Nearly all of them; even permissive BSD-like licenses usually(?) require that much.
For the compiler bug, as someone who never strayed into those regions of Rust programming, it's not clear to me how likely it is that someone would write such code by accident and introduce a memory safety issue into their program. I browsed the Github issues, but couldn't find an indicator how people encountered this in the first place.
The documentation addresses this case specifically: https://doc.rust-lang.org/stable/std/os/unix/io/index.html#p...
"Rust’s safety guarantees only cover what the program itself can do, and not what entities outside the program can do to it. /proc/self/mem is considered to be such an external entity..."
That doesn’t mean that it never happens, but it is a very rare occurrence, not something that Rust programmers deal with regularly.
> The author has absolutely no fucking clue what the code in this project does. It might just fucking work or not, there is no third option.
lol
&'a &'b
and not with where 'a: 'b
I suppose I don't understand enough of how the first is different from the second one, and how it impacts this issue.Also, it's a lot less weird if you don't stop in the middle of the type. The entire type is for example
&'a &'b u8
or &'a &'b ()
So the outer reference has lifetime 'a and the inner reference has lifetime 'b.Yes it does; there's an implied `'b: 'a` outlives relationship that's required for the type to be well-formed.
In contrast, if a lifetime is subject to an explicit where bound, it must be "early-bound": each function pointer can must choose one particular value for that lifetime, that must be upheld for every call. For a practical example, you might have tried to declare a closure with a &T parameter, only to get lifetime errors when you call it twice. This is because the lifetime in the closure type is early-bound and must be the same for every call. (Sometimes the compiler figures out that a late-bound lifetime is desired, but the rules are very subtle. This is also why it's difficult to write a function that accepts a generic async closure.)
Late binding is necessary for certain kinds of variance, which is a big part of this issue. Function pointers are contravariant in their parameters, which means if you have a type like "fn(&'short str)", you can cast it to "fn(&'static str)", since a 'static reference will always be valid for 'short. They're also covariant in their return type, so that "fn() -> &'static str" can be cast into "fn() -> &'short str". But when performing these variance transformations with late-bound lifetimes, the compiler doesn't always take into account the implied bounds in the source type properly, which allows you to perform casts that aren't actually sound.
where 'b: 'a
meaning 'b is at least as useful as 'a. So 'b can live the same time, or longer, but not less.