Rust RFC 2094: non-lexical lifetimes
github.com
github.com
And an eli5 https://www.reddit.com/r/rust/comments/6brtsu/eli5_nonlexica...
Variables exist in memory, right? But we only need them for some amount of time - then we want our memory back. There are many strategies for this - Java uses a garbage collector to search through memory and reclaim parts no longer in use, C programmers manually call 'free'. Java suffers from its strategy because of the performance impact of that scan/ freeing strategy. C suffers from the complexity of manually managing memory leading to many security vulnerabilities.
Rust takes this concept of how long memory should live - when it should be freed - and encodes it into its type system. Similar to how you can say 'this thing is an int and it can be added to other ints' rust says 'this thing lives this long and it can reference other things that live this long, or shorter'.
Then the compiler can tell you "hey, you're trying to access something that might have been freed" the same way it would say "you're trying to add an integer to a string".
This means you don't have the performance impact of the GC but you also have a compiler watching your back, telling you when you may be doing something unsafe.
The algorithm used to determine how long variables live and when they can / can't be accessed is the subject of the RFC.
I really like the section of "Thinking in Scopes" in the Book that describes this:
https://doc.rust-lang.org/book/first-edition/lifetimes.html#...
You can think of lifetimes as hidden variables defined by the compiler that are in a sense bound to some block of code. I find this to be the best way to think of them personally.
should be '... or longer '.
The reference only remains valid if the referent hasn't been dropped in the meantime, so it can only reference things that outlive it.
So, if you borrow something and try to store it somewhere with a lifetime the compiler thinks is going to outlive the lifetime of the thing you borrowed, that's a compiler error, whereas in a reference counted language or a GCed language the original entity just keeps existing for as long as required.
This is usually because said entity is on the heap - Rust puts everything on the stack by default.
You can of course put things on the heap, and there is a selection of types available in the standard library which are I guess a bit like C++ smart pointers that let you have heap allocation, reference-counted semantics etc. if you want it.
a = {}
b = {}
l = []
for _ in range(100):
if input('thingy'):
l.append(a)
else:
l.append(b)
In that snippet it's impossible to know how many references there are to a or b.In Rust, lifetime management is set up so that your code can statically determine when something no longer can be referred to (hence "free GC").
For example:
def f():
a = 3
g(a)
# a is no longer used from this point
return 5
you can clearly define when things are no longer usable, hence free-able. Python in this context will also free at the end of the function, but it's because it will be decrementing the reference counter and it gets to 0. But Rust won't need to have the reference counter since it knows it can free the object at any point after g.(This isn't a great example for a lot of reasons, Rust allows you to go much further than this)
This ends up being somewhat similar to writing code in more statically typed languages/languages with dependent types. You might have to rework your code so that the machine can properly identify when it can free memory.
This is because reference counting is prescriptive, that is, the count determines how long a value stays alive. Lifetimes are descriptive, that is, you cannot magically make something live longer by taking an additional reference to it, instead, you'll get an error.
It's not really like that - with reference counting you sort of have multiple 'owners' of a variable, and they all get to 'free' that memory except that 'free' just decrements a counter.
With rust there is a very clear owner at all points - you always know that your code in one spot owns the memory, and no other code does. There may be references handed out but those are very clearly distinct - that is, by looking at the code you can see "yes, I own this" or "yes, I'm borrowing this".
Now, with borrowing it's a little bit more like reference counting, but it's really just scope based. We know that the references are valid for some scope, and after that scope they are not. There's no internal counting or anything like that in the compiler as far as I am aware.
Actually there is another way which you'll find commonly in well written robust libraries - make the argument a const pointer, so that the callee is obviously not responsible for clearing it up and then the callee takes a copy to keep. This also allows passing both const and non const data to the same function, and means that there is much less likely to be an ownership screq-up across a library boundary. As a pattern though it can result in a lot of unnecessary copies, which Rust's system allows you to avoid.
So there is a formal sense of "ownership" of memory in rust? And right now it's based on the lexical scope in which the memory was allocated?
The language generally does specify that they are, and that's exactly how we teach borrowing today. This new one is not more specific, but more general: it allows all programs today to stay the same, but enables more programs that don't compile today to compile in the future.
For example, all non-immutable objects in garbage-collected languages (or similar features such as shared_ptr in C++) are incompatible with these restrictions. As another example, mutable global variables are generally incompatible with the semantics that the borrow check enforces.
That said, I could see analogous dynamic systems (rather than static ones as in the case of Rust) becoming more widespread. In fact, Transferables [1] in JavaScript are an example of one such system.
[1]: https://developer.mozilla.org/en-US/docs/Web/API/Transferabl...
Really interesting point about Transferables, I hadn't seen that.
So if it's successful we should see more of it but...we won't necessarily see many languages where it is as all-pervading as Rust.
In swift's case I'd think they would avoid too much complexity since the lifetime stuff won't be as pervasive as it is in Rust (so it can have more rough corners)
In a way it is great the work being done by the Rust team, as many companies wouldn't bother with it.
It is already influencing the design decisions of other languages.
However it really needs to be more ergonomic, I for example have been bitten trying to call a method into self from inside a closure.
From logical point of view it was clear, the closure wouldn't outlive the object that owned it, but the borrow checker thought otherwise.
This is one of the use cases being solved by non-lexical lifetimes.
It's not clear to me how far this goes. One annoying Rust idiom is having to allocate something, then pass it into a function so the function can return it. There's no way to give a return value the lifetime of the caller, not the callee. Will this change allow that? It doesn't look like it.
This is more in the other direction. "In the new proposal, the lifetime of a reference lasts only for those portions of the function in which the reference may later be used (where the reference is live, in compiler speak)." In modern compilers, when a local variable has been used for the last time in its scope, it's dead, and its storage can be reused. This is mostly done to free up registers. It seems to be exposing liveness analysis at the source level.
The use cases aren't that convincing. This may be more of a "because we can" feature.
What I see is actually "it's wanted/needed enough that it's worth putting in the huge effort to make such a complex (and novel?) system reality".
EDIT: to be more clear: what's new here isn't "exposing liveness at the source level" but "model liveness in region typing". If we were changing when the destructor runs (which we can't because backwards compatibility - also, doubtful we'd want that at all), that would be more "source level".
That is, in the general case, not physically possible, and this sort of thing only works in GC languages because they return GC pointers to dynamically allocated data. It certainly doesn't exist in C or C++ (unless I completely misunderstood what semantics you wanted).
AFAIK there is no implementation of anything (not even JITs) which can allocate on the caller's stack. The closest anything gets is Forth where the call stack and data stack are separate and returning just leaves data on the stack for caller to read.
Without a split stack, returning something larger than a register (or two) is done by passing a pointer to space on the caller's stack, to the callee.
There have been very specific schemes proposed for `-> [T]` in Rust, e.g. the slice is left of the callee's stack, and a pointer to it returned, so the caller can allocate that much stack space and memmove'd it up.
But that wouldn't help with "allocating" on the caller's caller's stack because moving the data would invalidate references to it (and if you can have a reference to something, you can use it in ways the compiler can't trace it, unlike a GC).
y = f(g(h(x)));
is really var hreturn;
h(x, hreturn);
var greturn;
g(hreturn, greturn);
// hreturn is no longer live
var y;
h(greturn, y);
// greturn is no longer live
Looked at this way, a function can return any fixed-size type without a copy. Some languages, mostly those in the Pascal/Modula family, did this
explicitly. Instead of having a return statement, within a function, the name of the function was the return value, and code could assign to it, or pass it by reference to another function.And Rust does RVO, in fact it's important for "placement new" structures until the dedicated syntax lands.
How is this similar to anything JavaScript does?
I give up on this thread. :) I can code in both languages, by the way, and I could swear that Rust is just a systems-level version of JavaScript, or JS before the hipsters got their hands on it and made it insufferable.
It might just be me, I wouldn't take my off the cuff commentary too seriously. I was kind of surprised when I started to learn Rust that the Internet makes it sound like this incredibly difficult and arcane language, and I thought, after getting a grasp on the basics, that it was much like a fine-grained Node. It's a nice language, I like it, personally.
Source: I was involved in Rust's design from the very beginning.
I'll take that as canonical, then. I'm probably showing my naivety, but I don't grasp the reaction to the comparison. Both are basically C-family languages, there is a distinct Mozilla connection, and recent Rust development does seem to be web-focused. Rust has a lot of features that one would like JS to have, or get right, if it does have something comparable.
I think I'm missing some details. Is there a back story for why Rust developers wouldn't care for that comparison?
And in particular, this thread is about static analysis of pointers to improve ergonomics while retaining memory safety. Javascript's approach to memory safety is just to use a garbage collector, which is essentially the opposite of the approach taken here.