Pre-Pooping Your Pants with Rust
cglab.ca
cglab.ca
Delaying 1.0 by a few weeks may seem like a big deal, but ultimately it is a self-imposed deadline. It's great to have those, but following them dogmatically might not be the best strategy. In cases like this, I generally lean towards slowing down and doing things right: otherwise you will pay the price ten times over later.
(That said, while I understand this issue, I don't know very much about the context in the Rust community, so I'm not actually sure that mem::forget should be unsafe. It was just the impression from the article and from previous Rust code I've read/written.)
[1] http://www.informit.com/articles/article.aspx?p=1216151&seqN...
I developed a reference-counting system in my library "upb" that handles cycles and never leaks objects. It is precise for no-longer-mutable objects (objects which may no longer change the set of other refcounted objects pointed to) and less precise for mutable objects.
The basic idea is to reference-count groups of objects instead of single objects, and define the groups such that no reference cycle spans groups.
If you introduce a link between two refcounted objects A -> B, then A and B's groups are merged. Now refs/unrefs of A or B ref/unref the (merged) group. Nothing in the group is freed until both A and B's refcounts both fall to zero. This conservative group-merging whenever you create a link ensures that no reference cycle spans groups.
This is imprecise and relatively wasteful if the group grows really large. But you can make it totally precise for any subgraph of refcounted objects that you're willing to freeze. At freeze time, compute strongly-connected components, and each SCC becomes a refcounted group. For any frozen subgraphs, the refcounting is totally precise. And unfrozen objects can reference frozen ones (but not the other way around).
If your application ends up freezing most of the graph, this works great. Or if you're ok with the collection being pretty coarse for your groups of objects, it also works great. If you have an application that can't freeze the graph and wants pretty precise collection, this doesn't work so well.
More info about my scheme is in comments in these headers:
https://github.com/haberman/upb/blob/master/upb/refcounted.h
https://github.com/haberman/upb/blob/master/upb/refcounted.c
I always was curious if the semantics of this scheme would play nicely with the ownership system in Rust.
The minimized groups would still merge with other groups if you created links between them. But this would be a way to make the refcounting more precise without strictly requiring that subgraphs become immutable.
For project like Rust deadline doesn't actually matter that much, grasping for it is almost stupid. The only thing that actually can suffer from breaking it is self-esteem, which should't worry a reasonable man much. But magic "1.0" number does matter a bit more than just deadline, because after it there's no "breaking changes".
So I'd be much happier if Rust wouldn't reach 1.0 for the next 2 years, but would actually become satisfying instead. Somehow "non-stable but working" is better suited for making software than "broken and stable". We have plenty of "broken and stable" out there already.
For example, Java has the function System.runFinalizersOnExit, but if you look at the documentation you'll see that it now says, "Deprecated. This method is inherently unsafe. It may result in finalizers being called on live objects while other threads are concurrently manipulating those objects, resulting in erratic behavior or deadlock." http://docs.oracle.com/javase/7/docs/api/java/lang/Runtime.h...
The problem is discussed in detail in Hans Boehm's 2002 technical report “Destructors, Finalizers, and Synchronization”. http://www.hpl.hp.com/techreports/2002/HPL-2002-335.pdf
Boehm's key points are: (1) When a destructor on an object O runs normally, anything that O points to is still alive (because of the reference from O), but when a cycle is collected, not everyone can go first: that is, all but one destructors in the cycle have to run after one or more of their references are dead. It is hard to write destructor code that is safe in all cases. (2) Any destructor that needs to update a concurrently accessed data structure has to take a lock, but destructors on cycles run asynchronously with respect to the rest of the program and so an unlucky timing leads to deadlock.
If this problem was better appreciated then language designers wouldn't go down the rabbit hole of trying to figure out how to make destructors work together with automatic collection of reference cycles, and instead try the alternative approach of providing mechanisms for the program to run destructors synchronously.
In the Memory Pool System we use a message-passing interface: http://www.ravenbrook.com/project/mps/master/manual/html/top...
So you cannot safely assume your destructors will be called, and you need to proactively take steps like this in case they aren't. Nothing in the natural use of the language cues you to do this, so you need a moral guide like "Effective C++" for Rust.
This was the kind of accidental complexity that Rust had hoped to avoid. It seems like Rust has avoided a lot of it. But I guess it was entirely too optimistic to hope Rust could avoid all of it.
Edit: I was wrong. The problem isn't with forget. This is hairy.
Which, as you can see, uses Rc, which uses unsafe code in a way that leads to the soundness bug. So that's not strictly true, or rather, doesn't really change anything, as it relies on the same bug.
"1. There were other ways to forget even before Rc, at least in some cases. For example, if T:Send holds, you could send the value to a thread that runs an infinite loop, or which is deadlocked on a port.
"2. It is true that one cannot assume that a destructor will run, and hence that forget is not itself unsafe (rather, it is unsafe to write a dtor that must run)...."
And alexcrichton:
"I commented on #24292, but the gist is that there are multiple ways to leak memory today (e.g. #14875 and #16135), so a targeted solution at Rc may not cover all use cases. Although as I mention in #24292 these other bugs can also be considered separate bugs on their own which need to be fixed regardless (but sometimes is quite difficult to do so)."
I especially liked dgrunwald's comment:
"A safe mem::forget has the advantage that it makes it easier to write the counterexample proving thread::JoinGuard unsafe. Safe mem::forget makes it more likely that people will know that destructors are not realiable, so they can avoid repeating this mistake."
But then, "Put a big freaking spike in the middle of the steering wheel and get rid of the airbags and seat belts" has always appealed to me as an automotive safety approach.
But the unsoundness RC bug still relies on unsafe code which was written incorrectly.
When the footgun goes off in unsafe code in the standard library, you get use-after-free and memory unsoundness. When the footgun goes off in safe code written by mere humans, you leak arbitrary resources.
The fact remains that as long as you are not using unsafe code yourself, you don't really have to worry about anything. The category of destructors that leave the world in an unsafe state if not run is AFAIK a strict subset of the category of destructors that require unsafe code to be written.
And Rc only looks like it is the problem. Rc is doing exactly what it says it does, as safely as it can. The fact that you can use it to avoid having a destructor called is a side-effect of it doing what it's supposed to do.
If main is safe, mem::forget is safe (remember, Rc is just playing the part of mem::forget), and the destructor for JoinGuard is safe (which it is; I just checked), where is the unsafe code?
(I myself have been guilty of this same error, of assuming that `unsafe` regardless of its position cannot have far-reaching effects on the overall state of your code.)
The important thing is this: using safe Rust, and using any safe stdlib API, memory unsafety is impossible. This includes stdlib APIs that use `unsafe` internally, such as `Rc`. Proving those interfaces safe is the burden of the Rust developers, and if you can show that there has been an error in these proofs then it is a drop-everything sky-is-falling defcon-1 situation that they must address, up to and including unreleasing the language if no suitable solution came to mind (fortunately, this is quite unlikely).
In the thread spawning. The destructor itself isn't unsafe, but it was being used as a guard to enforce the safety of the threading code.
> I can easily imagine another situation where a programmer assuming a safe destructor is always called causes a bug, which is exactly the problem with thread::scoped.
There's a difference between "a bug" and "memory unsafety". If the bug is that the value the destructor guards isn't released (for example, if a File is leaked, the associated file descriptor will never be flushed and closed), then that's not Rust's problem. The programmer probably wanted that file descriptor to be closed, but there's no memory safety violations going on there, and in fact nothing worse than simply stuffing that File into a global array/hashtable and forgetting about it.
The problem with thread::scoped is actually fairly unique. The destructor is not doing much besides just calling join() on the forked thread. The problem is that the JoinGuard object doesn't actually control any access to anything at all, it exists purely as a proxy to represent any values (particularly stack values) that are borrowed by the forked thread. Because it's only a proxy, leaking the guard does not actually make the guarded values inaccessible.
I can't think of any situation besides multithreading that has this same problem. Every other situation involving a RAII object with a lifetime being leaked (and therefore outliving its lifetime, which is technically a violation of borrowck) by definition makes the object inaccessible, which means the data it's protecting cannot be reached. For example, a struct that contains a &-ref to some value may outlive the referenced value if the struct is leaked, but this is still memory-safe because the reference will never be dereferenced. If the reference can be dereferenced, that means code exists that has visibility on the struct, which means the struct hasn't actually been leaked yet, and as long as it hasn't been leaked it's still safe (e.g. because borrowck will still enforce the lifetimes).
Multithreading is special because the forked thread is what's actually accessing the values, without ever going through the RAII object. This means that borrowck cannot guarantee that the RAII object is still alive (and therefore that the "protected" values are still valid). This was considered to be ok as the RAII object would join the thread (therefore ensuring it had finished executing) before destructing, but the ability to leak objects means the join may not happen.
Ultimately, this means that JoinGuard is the moral equivalent of a Mutex that coexists with its protected data rather than owning it (thereby making it possible to write code that accesses the data without holding the lock). It's a bit better than that, because you have to actually go to some effort to leak the JoinGuard as opposed to just doing it accidentally, but the point is that borrowck cannot guarantee that the code that accesses the data does not outlive the guard.
And all of this comes back to my original point, which is that I can't think of another situation where it's actually unsafe to leak the RAII guard (as opposed to merely being a bug) without there being `unsafe` code in the guard itself. Note that the example given for PPYP, Vec::drain_range, uses `unsafe` code in its destructor.
Sometime programming can be boring and then you come across PPYP and it makes it fun again.
I don't think it's always the same people or that it has anything to do with Ruby itself, it just seems like there's a certain culture that first grew in the Ruby / Rails community.
> I like the naming and feeling of "Pre-Pooping Your Pants" pattern. We are doing the intentional leak only to be (hopefully) collected later, which may not happen but it still remains memory-safe even in that case. Still this is not very desirable and the name clearly indicates so.
Not all hacks are bad. But, if I see "// this is a hack to..." or "# this is a hack to..." or similar in code, that is a signal to me that the developer (1) either didn't have time to perform the right thing but could have with more time (within their lifetime), or (2) could not feasibly perform the right thing under constraints given (within their lifetime).
Over the past ~8 years, "hack"'s use has switched back to a much more positive sense- possibly because of HN, or possibly because people read Eric Raymond http://www.catb.org/esr/faqs/hacker-howto.html, but "hack" != "hacker". The negative sense is still used and understood by most developers, and it comes from two things: "HACK... HACK... HACK...", like the sound of an ax attempting to chop down a tree (brute force), and the other definition of a hack not used as much these days, which is "work for hire especially with mediocre professional standards".
I think we should at least consider PPYP to be a subclass of "hack", or perhaps just use the word "hack" instead of PPYP.
How am I going to show this to my boss, and have him take it seriously with such an infantile title?
And part of that is showing flaws, and how they're resolved/worked around and so on.
EDIT: I do not want to argue with Dewey2, so I'm putting this here: the women in the rust community that feel uncomfortable when they are addressed as "guys" have every right to feel that way, and I commend the rust community on making that stand.
Probably the same way you would explain to your boss what KISS principle means or what WTF-metric for source code is :)
Illustration: https://davidlongstreet.files.wordpress.com/2009/04/wtfm.jpg...
If the author had written this in a modern, safe language like Rust this would not have happened.
But the problem here is complex and hard to solve. Cycles and memory leaks are nigh impossible to solve completely without sacrificing on that altar ability to create custom data structures or no-GC by default. Which was non-starter for Rust.
After all, why Rust? Because we're all tired of problems with C++ and stuff. And the point isn't really about making a language that would be safe "as often as possible", it would be quite terrible goal, actually. The point is, I believe, to make it nearly impossible to write unsafe code accidentally, without even noticing it.
And again, maybe it's just me, but what I understood from that post is totally non-intuitive to me. I don't really understand what are the guidelines and how many more gotchas like that there could be. "Not to mess with unsafe code" is not really good guideline, because it turns out that it can be not quite obvious that the code is unsafe.
I assume much of it is safe, with few probably hidden caveats, hidden in more complex logic. In C++ writing code that causes this behavior is trivial, here you need to jump through a lot of hoops. So it offers you comparatively more safety, but this kind of bugs are potential gold mine for crackers.
Anyway core developers are on the scene deciding how to tackle this issue.
> but this kind of bugs are potential gold mine for crackers
Exactly! And I'm just saying that a problem which is "not very likely to trigger, but is very hard to find if happens" is far, far more dangerous than a problem which is "easy for a novice to miss, but every experienced developer would notice".
> If things are complicated enough that we actually need Agda/CoQ/Idris to
feel safe — seems that things went really wrong at some point.
I disagree. Complex software/hardware is by its nature full of bugs. What Rust set out to do is complex and it's bound to have bugs, but doesn't mean it's reason to abandon it, no more than we should discard LHC or ITER because they had issues starting up.Whether it's a wrong invariant on TimSort or a really obscure multi-threading case, or a weird glitch in time library, its still an error. Each one is exploitable.
ALL software could use some coverage/tests and theorem proving. I just wish Rust started with a proven safe subset.
> Exactly! And I'm just saying that a problem which is "not very likely to
trigger, but is very hard to find if happens" is far,
far more dangerous than a problem which is "easy for a
novice to miss, but every experienced developer would notice".
Again, I disagree. The matter of security is one about making the assailant less likely to breach your software. If you take two comparable pieces of code, the one with more holes will take less effort to break.If your language offers fewer avenues of exploit, makes these exploits more expensive. A rare bug is worth more than a run of the mill bug. Which prohibits the number of assailants which could purchase it, which reduces number of potential attackers.
( people getting mad about PPYP is really a victory, though; the second last bullet was the most important:
> It's vaguely incoherent and meaningless on its own: this is a property inherited from its connection to the venerable Resource Acquisition Is Initialization (RAII) pattern.
I'm mad about RAII and not going to take it anymore! )
I too am particularly terrified of this kind of leak. The reason is that it is the only kind of leak mentioned that is not clearly a programmer mistake.
C++ has a whole class of "mistakes" that one cannot be aware of simply by reading the language specification. Common ones such as writing a constructor and destructor but not a copy-constructor are now compiler warnings, but many will pass under the radar. As a result, to write correct code one needs to be aware of multiple conventions and rules and patterns that are external to the language. When reviewing code one needs to do more than observe that it compiles and accomplishes its stated task, there are probably several organizational standards that need to be satisfied as well to be safe.
This is precisely the kind of thing that Rust was supposed to solve. The goal was to have a language that didn't require an "Effective <xyz>" book to be memorized to do code review. A language where any rule like "You should always free resources in advance in case the destructor doesn't run" is captured statically by a compiler error.
So yes, I am particularly put off by this error.
The guarantee that destructors are called when values leave scope is fantastic. It means that one can leverage memory-safety into arbitrary-resource-acquisition-safety. It violates the principle of least astonishment in a brutal way if Rust's borrow checker can statically guarantee that a value no longer exists, but can't know whether or not its destructor has been called.
This only shows up if you're specifically opting into unsafe code by using the "unsafe" keyword. Rust doesn't (yet) try to ensure safety of unsafe code, because it's unsafe. As long as you don't step into unsafe code, you can ignore this entire article.
[1] Which creates a reference-counted cycle containing the thing-to-be-forgotten, ensuring the finalizer is never called.
[2] The first main simply demonstrates that the destructor is never called, which is ok in itself but causes problems when you're doing smart things in the finalizers, as in the second main, where the code uses the destructors to ensure the threads are terminated and the programmer is assuming the borrow checker is going to warn them if the threads aren't terminated.
But the use of Rc is kind of a red herring; if mem::forget is marked safe, then you don't need safe_forget because you can forget things (fail to execute their destructors, specifically) safely anyway.
The problem is that destructors aren't guaranteed; the bugs in the thready thing (and potentially unbounded other things) are symptoms. Drop needs big, red warning signs.
The power of a notation, and a type system, is in what it lets you not think about. The fact that destructors may not be called is, unfortunately, something you have to think about.
As evidenced by this whole kerfuffle, you can in fact leak heap-allocated values by constructing an Rc cycle (or taking advantage of various implementation bugs in other things), or by using mem::forget() (which I believe is still marked as unsafe but seems likely to change). But none of these approaches could be characterized as "forget[ting] to ever free it", because they're not a consequence of "oops I forgot to do that".
Most important, this API problem does not break Rust's basic safety guarantee. That is, you must write unsafe code to violate memory safety.
If Rust's basic memory safety guarantee were at risk, the core team would absolutely consider delaying the release, or doing whatever else was needed to address it! We must never allow memory safety violations in safe code, no matter how obscure the bug that leads to them.
So this is a question of writing "safe" APIs that hide uses of `unsafe`. What precise assumptions are you allowed to make within such code?
As it is, we are taking this API issue very seriously, and making sure that we explore all of the options available. We have already yanked the affected APIs, while we determine the best way forward.
There are known ways to safely re-introduce the (relatively few) places in the standard library where unsafe code was using this pattern. Gankro talks about one in the post; you can see a proposal here (https://github.com/rust-lang/rfcs/pull/1084) for the `scoped` API.
The basic issue here is about a tension between a couple of different APIs, as Gankro explained in the post. In particular:
- `Rc` has been part of Rust for a long, long time, and is a fundamental systems programming tool.
- APIs like `scoped` and `drain_range` would like to use RAII for ergonomics and consistency, but this is either not possible (`scoped`) or subtle (`drain_range`) given `Rc`. In either case, these APIs use `unsafe` internally, so it's all about what that unsafe code can assume.
Furthermore, there is not a strong consensus about how best to resolve these issues, even if we had all the time in the world. Currently there are at least three camps:
- Stick with the status quo (perhaps marking `mem::forget` safe to reflect that reality). Safe code can assume no leakage, but unsafe code needs to take extra care (and sometimes avoid the RAII pattern). This is already the case for a large number of other properties, as Gankro points out. Leakage in safe code can never lead to memory unsafety; it is just a vanilla bug, and Rust doesn't prevent you from making bugs.
- Introduce something like `Leak`, meaning that you have to explicitly ask for a given type to be guaranteed not to leak through things like `Rc` cycles. While that allows you to write APIs like `scoped` and `drain_range` easily using RAII, you need to be aware of the marker, and make sure to use it for such types. Worse, though, is the interaction with trait objects: depending on the design, you may have to write `Box<MyTrait + Leak>` to be able to store the trait object behind an `Rc`, and that `Leak` bound needs to be present for the entire chain of APIs leading up to that point.
- Restrict `Rc` in some way, perhaps to `'static` data, thereby ruling out (a class of) leaks for all types. It's not completely clear how much fallout this would involve. The compiler currently relies on non-`'static` reference cells, and this change would likely force channels to use a `'static` bound as well, thereby defeating much of the purpose of the `scoped` API in the first place.
Finally, it's worth noting that at least one version of the `Leak` proposal can be added backwards-compatibly, later on, so there is potentially plenty of time to explore that approach if we feel the complexity is worth it.
I believe that Niko Matsakis (also on the core team) is planning to write a blog post explaining all of the above in much greater detail.
impl<'a> Drop for Foo<'a> {
fn drop(&mut self) {
*self.0 += 1;
}
}
What does the line "self.0 += 1;" do? Specifically, the ".0"? Foo is a basically a reference to an i32, so I would have thought that would be "self += 1;".