Rust 1.0: Status report and final timeline
blog.rust-lang.org
blog.rust-lang.org
One thing I'm not clear on it if can do, and that I'm interested in, is secure destructors.
Say I'm handling crypto, and I'm carting around an ephemeral key. When this goes out of scope, I definitely no matter what, want this zeroised by its destructor - as opposed to just having it (or a temporary copy made by a compiler optimisation!) zombling around the heap, stack or forgotten unused xmm registers because the compiler figured since I don't reference it again, the memory's contents are no longer important.
Current approaches to this involve explicit_bzero(), or other similar memset(0)-and-I-really-mean-it-don't-optimise-this-out techniques. (And a fair bit of testing and prayer when it comes to potential temporary copies or registers.) But unless you're doing it in assembly language, you don't really know. (The stack beneath you, such as the OS, any hypervisors, SMM, AMT, SGX, µcode etc, aside, of course!)
I'm not quite clear what Rust's behaviour with this scenario is. If it can do this easily, even potentially, I am very interested…?
But what I'm a bit more worried about is how the secret data gets in there, and what happens while I'm working with it: expansion, cipher state, key setup, all the little adds and xors and rots (dammit, why doesn't ROT ever get some real operator love? It's got first-class instructions… :() - all that stuff you'd do in u32 and u64. Temporary copies may still be a problem, if you look at the object that actually comes out of compilers sometimes.
Does using an opaque type actually deal with that issue here?
Unless you can actually prove all parts of the compiler's data transforms going down to assembly I think the safest thing to do is sandbox your key-handling process so nobody else can examine it.
Yes, actually proving all parts of the compiler's data transforms going down to assembly kind of is what I'm after, if we can get it…
There's a lot of interest in our community to build out a solid foundation and verified foundation of cryptography primitives. If you're interested in helping out, hop in the #rust-crypto channel on irc.mozilla.org.
[0]: http://www.meetup.com/Rust-Bay-Area/events/210632582/
[1]: https://air.mozilla.org/bay-area-rust-meetup-december-2014/
If you are running in an environment where you don't trust code running in the same compartment/sandbox/process, then it's futile to zero out memory. The caller could have prepared things such that the memset doesn't work, if the key material went somewhere else.
If you ever find yourself thinking you need to do this, what you instead need is a helper process who's only purpose is to do primitive operations with sensitive key material.
Particularly as Rust is already a "safe" language - it doesn't even make sense to zero memory which by definition another piece of code can't access. Unless there's declared "unsafe" code lying around, but you wouldn't put that in the same process, would you? At which point, what are you even protecting against? If an in-process threat is that advanced, then you're not achieving anything.
Maybe zeroing destructors make sense as defense-in-depth, but I don't see how they can fix a Heartbleed-style exploit in Rust. In code where the buffer is freed and its destructor runs, Rust's memory safety guarantees already prevent it from being accessed after free. In vulnerable code that just uses the same buffer twice, the destructor never has a chance to run so its behavior doesn't matter.
The real Heartbleed vulnerability (CVE-2014-0160 in OpenSSL) involved reading into uninitialized memory in a newly-allocated buffer, which safe Rust code already prevents [2].
The point is Rust already provides safety guarantees. If you don't trust the runtime, then why would you trust the built-in zero'ing? I get the "defense in depth" argument, but it feels a bit like doing this:
{
int a = secret; // Get secret.
assert(a == secret); // Check "a" is actually that.
a = 0; // Ensure "a" is zero'd on exit.
assert(a == 0); // Just because.
}
And yes, I get that you can build this into the language so it's not quite as ridiculous - you actually wipe tainted stack, for example.But the point is: the runtime has an ABI and a machine model. Information is allowed to leak across function boundaries, beacuse it doesn't matter. Without using the "unsafe" keyword, there are no methods of getting around the machine model and dipping into the underlying actual machine.
Even if you don't have a "safe" language and runtime, it's still of limited value. It protects against threats involving data or control flow corruption after key usage, and where there isn't sufficient control of the program to perturb the secret-consuming functions. That's more of an annoyance than prevention. On the other hand, it gave the programmer a false sense that it was properly wiping secrets.
We actually have an interesting project in rust where someone is writing a syntax extension to take rust like code and generate assembly [0]. It's probably unsafe to use right now but if sufficiently well implemented it could be the foundation of a lot of interesting cryptography work.
There are many ways to dig out stale memory if you're running at sufficient privilege, for example. Direct cache introspection, for example, or bypass. Zero-izing alone is not sufficiently strong to mitigate the threats people imagine it works against.
Zeroing out memory after use is mitigation against security flaws like heart bleed. It's not a security feature in it self. Although you are right that the secure process is probably just the better solution anyway.
It's sort of like the arguments for DRM. Of course pirates will always break it, but we can still mitigate it in a number of ways.
It's the same idea as in "forward secrecy": once the key material is discarded, it's gone, and no future compromise can bring it back. Being able to say "after this point in time these values don't exist anymore" is a powerful cryptographic primitive.
I want to be able to have an ephemeral secret, and then trust the code to do its very best to get rid of it when it's no longer needed. That doesn't mean just leave it rotting on the heap and promising it doesn't get accesed again, it means burning it to make sure. That's the underpinning of any possible proof of forward security.
Sure, you say, I want a helper process? Fine idea: compartmentation. Now that helper process needs secure zeroisation. And since I want it to be secure, surely I want to write that process in Rust. See where I'm going here?
Whether 'safe' code I trust within my environment can read it again is totally irrelevant, if the machine later gets rooted or booted, nonstopped or whatever.
Of course this would only be possible for languages like C which don't make any advanced abstractions on the hardware arch.
Python's "with" clause seems to do the best job of making sure closeout is handled properly. Take a careful look at the extra arguments to a __exit__ function in Python,s "with". It's one of the few closeout methods where things such as an exception during closeout (an I/O error during file close, for example) is handled in a way that doesn't interfere with other closeouts.
I've been building a 3d game with Rust and OpenGL, ported from a C++ codebase. So far, my experience has been very positive. Despite Rust's supposed immaturity, it feels more polished than C++ in many ways. Forward progress has been much faster than it was with C++.
Does anyone else have a story (positive or negative) about using Rust in real projects?
The challenge with F# is controlling memory usage. Even one extra allocation per packet can make a measureable difference in performance. I ended up doing a ton of unsafe code and manually managing most of the heap. Rust allows me to write fairly high-level code (not as expressive as F# yet but whatever) while getting "best" performance. Inline asm is a bonus, as there's some algorithms for integer compression that can use SIMD for big wins (I can do that in .NET, but it's ugly, and doing it safely means a ~30 instruction thunk). And sometimes in tight loops, I've found it difficult to get .NET to do acceptable codegen, causing double-digit% impacts.
There's also the safety issues writing unsafe code for network-exposed traffic. So Rust is actually more safe than .NET, because I have to toss .NET's safety to gain performance.
The backend management code I can continue to write in F#, and Rust's C-compatibility means it's trivial to interop the code. So I can do "orchestration" of indexing daemons and management APIs and such things in a higher-level language, then for actual indexing and whatnot, just jump over to Rust, seamlessly.
Finally the static compilation means a smoother installation experience for customers. And if I ever ship a closed-source module that executes on the client, I don't need to license Mono for static linking. So that's nice. And the safety guarantees are good, because similar, existing, software in C has put customers at risk before. (I'm not sure if I can effectively market that last part, but hey.)
Rust would appear to have a unique value proposition and I'm very pleased to see it progressing so damn well.
The alternatives you listed aren't known for being able to write top-performance idiomatic code (I've got something _working_ in F#, but it's ugly non-idiotmatic code). The overhead of a GC is just too much to pay when doing linerate networking. Rust allows me to keep nice, high-level, idiomatic style, without paying any overhead. I can account for almost every byte.
I hate to let such a triviality lower my enthusiasm for a language so much, but I just can not get over that awful inconsistent closure syntax :/
I don't get it. Most everything else has a nice unique keyword syntax, fn uses (args, in, parenthesis), proc syntax made consistent sense, then lambda is this crazy || linenoise thing that doesn't fit in at all. The "borrow the good ideas from other languages" approach has resulted in a great language, but "cram random syntax from other languages that doesn't fit" doesn't work out so well.
x = foo {|x| x + 1 } # Ruby
let x = foo(|x| { x + 1 }); //Rust
let x = foo(|x| x + 1 ); // single expressions don't need {}s
That said, I'm not sure what the exact reason was for choosing the syntax, as that was before my time.I think the reason is just that it's very concise, and lightweight closure syntax makes things like `Option::map` feel like first-class parts of the language. The closure you pass just sort of seamlessly "blends in".
Note that having especially sugary here is not so uncommon, for example Haskell has `\x -> blah` for Rust's `|x| blah`.
I'm personally very happy that the closure syntax is as concise as it is.
Hand-writing a parser for some other language leads to madness - just ask the folks who've done SWIG, GDB, or most IDE syntax-checkers. You'll inevitably get some corner-cases wrong, or the language definition will change underneath you long after you've ceased to maintain the tool. Instead, the language should just expose its compiler front-end as a library, and then you can either serialize the AST to some common format for analysis outside the language or build your tools directly on top of that library.
By making the language simple you can easily implement your own parser. This opens up the ability to write native parsers in other languages, say vimscript. By keeping it super simple there -are- no corner-cases.
There are many benefits to this (like the formatters etc that others have alluded to) from things like IDE integration (imagine lifetime elision visualisation, invalid move notifications, etc) static analysis tools and more. None of these tools then need to be written in Rust. It also means it's easier to implement support in pre-existing multi-language tools.
Don't underestimate the necessity of a simple parseable grammar. Besides, people have endured much worse slights in syntax (see here Erlang).
The ones I actually use all call out to the actual compiler - Python, Go, or Clang for C++.
Just because people write their own parsers doesn't make it a good idea. It may've been necessary when most compilers were proprietary and people didn't have an idea how to make a good API for a parser. But now - just don't do it. You'll save both you and your users a lot of pain.
> Closures: Rust now supports full capture-clause inference and has deprecated the temporary |:| notation, making closures much more ergonomic to use.
|args| expr // upvars captured by reference, can't be called after function has gone out of scope
move |args| expr // upvars moved from function to the closure context (or copied if trivially copyable)
This is simple and good enough for most use cases. If you want more complex schemes, you have to implement them manually, e.g. to reference count the upvars, like Apple blocks do by default, wrap them in Rc and capture that.Accepting closures is a bit more complicated though.
For a brief period, closures had to be annotated in certain cases like |&: args| or |&mut: args| or |: args| to determine whether they captured their environment by (mutable) reference or by value/move.
Now that this is inferred in all cases, closure arguments can just be written as |args| in all cases, just as they were before the current Fn* traits were introduced.
That particular annotation controlled the access a closure has to its environment, not how it's captured. |&:|, |&mut: |, and |:| corresponded to the Fn, FnMut, and FnOnce traits, respectively. If you look at the signatures of those traits, you'll see that Fn's method takes self by reference, FnMut takes it by mutable reference, and FnOnce takes it by value. In particular, this means that the body of an FnOnce closure can move values out from the closure (that's why it can only be called once), whereas Fn and FnMut can only access values in the closure by reference and mutable reference, respectively.
The way variables are captured from the environment into the closure is controlled by the "move" keyword. If the "move" keyword precedes a closure expression, then variables from the environment are moved into the closure, which takes ownership of them. "move" is usually associated with FnOnce closures, but it's also needed when returning a boxed Fn or FnMut closure from a function, as you can see below:
fn make_appender(x: String) -> Box<Fn(&str) -> String + 'static> {
Box::new(move |y|
// The closure has & access to its captured variable, but
// it has been moved into the closure so it outlives the
// body of make_appender, thanks to the move
// keyword.
x.clone() + y
)
}
fn main() {
let x = "foo".to_string();
let appender = make_appender(x);
println!("{} {}", appender("bar"), appender("baz"));
}Note: if you don't specify `move`, then the captures are determined in the usual way:
`|| v.len()` captures `v` via an immutable borrow
`|| v.push(0)` captures `v` via a mutable borrow
`|| v.into_iter()` captures `v` by moving it
Edit: as someone involved with other, less mature (and less ambitious) open source projects, if you know of pain points in the governance of rust, i'd be interested in learning about them.
This page talks about the process and their code of conduct: https://github.com/rust-lang/rust/wiki/Note-development-poli...
I was under the impression that a few primary contributors (mostly/all mozilla employees?) are gatekeepers to merging anything.
Having an "RFC" issue tracker isn't the same as having a governance structure.
Edit: I suppose you could call the above a 'governance structure', but I'm having a hard time seeing anything impressive/different about it from other open source projects
PRs can be reviewed by anyone from a large pool of reviewers. PRs that introduce new features etc. have to come after an RFC is approved, however.
The core team includes Huon Wilson, Yehuda Katz, and Steve Klabnik, none of whom work for Mozilla (though the latter two have done some contracting work). We hope to continue expanding to include other stakeholders.
EDIT: Steve tells me he's currently working as a "seasonal employee" at Mozilla for doing Rust docs, but it's a short term thing.
Maybe a third to a half are Mozilla employees (although it's infamously hard to tell who actually works at Mozilla and is just weirdly into maintaining Rust).
It looks like April will be a good month.
At least optimized builds aren't affected, but it sounds like lots of code (including Rust nightlies) aren't built optimized.
At the same time I don't think I would use Rust 1.0 in production for two reasons :
* the language doesn't seem to be mature enough to be highly productive, for eg. the type system isn't powerful enough to express stuff like Iterable or VectorTN and I'm sure there is plenty of tedious stuff like that along with pains from ownershinp systems
* tools and libs are obviously not there
So I guess I'll wait for early adopters to write the libs and give feedback on their painpoints to the devs.
I've said this before - I think Rust 1.0 is something that I could use (ie. working and stable) but I don't think it's something that I'd want to use yet.
It seems like the standard library has seen a ton of work in the last month or two. I'm surprised at how aggressive the release schedule is, given how things are still churning a good deal.
> For those who are currently more in touch with the situation, does this
> release date look realistic, without quality suffering? Is the 1.0
> release premature, aggressive, or very realistic?
I think it's aggressive, but I don't think it's unrealistic. Personally I had hoped for a late June/early July release to give more time to solidify the docs and to shake out bugs in the compiler.There will definitely be people who say that the release is premature, but I'm personally not one of them. The core of the language is ready, even if some pieces around the edges could still use some refinement (and will see refinement, backwards-compatibly, in the coming releases).
> I'm surprised at how aggressive the release schedule is,
> given how things are still churning a good deal.
You'd be surprised at how much of a motivator a concrete release date is. :) Churn is happening now because of the impending release, not despite it. The language intends to have a solid compatibility story (via semver) for post-1.0 releases, so everyone who'd been holding off on changes for the past few years has suddenly come out of the woodwork to implement them.Rust and Go are two reasonably new languages. I understand scope and priorities of each language may be different but my idea is to get some approximation/thumb rule for any one before starting similar journey.
As per Github and wikipedia:
Number of contributors for Rust and time taken so far: 840, and 2 years. Number of contributors for Go language and time taken so far: 424 and 6 years.
Financial details are not known.
It seems developing new language and bringing it to reasonable level is not trivial effort.
1. Is above data correct i.e. are those contributors full time working on those languages i.e. is it full time job of those people?
2. Can we get details like number of developers/number of test engineers/number of documentation writers ...etc?
3. Is it possible to know the total amount of financial resources consumed so far in the effort?
4. Is there any research into resources required for new language development in terms of man power, time, financial resources for various languages?
It is fascinating to see a new language developed in front of us.
In terms of volume of contributions, for Go, Google employees are by far the most active: https://github.com/golang/go/pulse. In this graph, the 7 top contributors to the project are Google employees. The Rust pulse graph shows the same trend (https://github.com/rust-lang/rust/pulse), with the top 6 contributors being Mozillians (according to a few Google searches).
Something that noteworthy about Go is the "quality" of the team members: Google has Ken Thompson, Rob Pike and Russ Cox working full time on the language. Mozilla may have a few great developers on Rust too, but Google is very serious about Go.
I don't have any information about how financial and human resources are used by Google and Mozilla for the development.
Rust and Go include the compiler, the runtime, some libraries, some package management and some tutorials.
Other things you might want include IDE support, static analysis tools, advanced garbage collectors, GUI toolkit (bindings), slick debugger support, monitoring and profiling engines. These will all add a large amount of time and money to the development of your new platform.
A scripting language or a JVM language or a PyPy based language (or pick two of those) can make it to a 1.0 much faster.
That is correct :)
Is there a "post 1.0 wishlist" somewhere?
We have not done any scheduling for what comes after 1.0. I've been doing a huge amount of triage to move over issues that used to be in rust-lang/rust to be under that label, too, and to make it more detailed than just 'wishlist.' We'll figure out exactly what's next shortly before the first post-1.0 cycle actually begins.
http://users.rust-lang.org/t/most-coveted-rust-features/324 was a thread created recently, asking people what they want the most. You might find that interesting too.
The thing about "efficient code reuse" is it probably requires dynamic dispatch. Once you have dynamic dispatch, you suddenly have vtables. But who decides what those vtables look like? Where do they reside in memory? What's the layout of that? If a struct suddenly has an is-a pointer, where is that mentioned in the code? Now my struct isn't just a struct.
I love the rust idea that tons of modern language design still allows for zero-cost abstraction. Inheritance in dispatch starts getting into the land of "putting a lot more stuff in my binary than I asked you to", and I would argue that it's this property more than anything else that keeps embedded programmers and kernel guys safely in the minimalistic land of C.
It would be nice to have a new C, finally. But the more a language has an opinion on runtime layout, behavior, and symbol names, the less C-like it becomes.
Rust got this right when it decided that GC was NOT the correct default behavior for a language. The reason you see people playing with OS kernels in rust and not as much in D is, I think, mainly due to this decision. I think rust should continue carrying this torch.
If not for the ability to have the best of both worlds (a modern language and access to to-the-metal programming with a controllable runtime layout and deterministic performance profile), where exactly is the value in learning how to use the borrow checker?
1: http://doc.rust-lang.org/book/static-and-dynamic-dispatch.ht...
2: http://www.reddit.com/r/rust/comments/2j78oh/i_heard_that_ru...
I know that's a hard ask, but the upside would be that anything in the language that's represented by a complex layout + behavior could in principle be replaced by another implementation that preserves the size and behavior contracts. If an implementation is not provided, those features of the language are unavailable. This kind of takes the "you can't use x and y until you give me an allocator" approach and turns it up to 11.
But you are correct in that one of the loudest objections to the potential inclusion of inheritance in Rust is that people don't want two ways of achieving dynamic dispatch, because suddenly then you're in the same boat as C++ where you must decide which incompatible subset of the language to use in your codebase.
1. We already have vtables through trait objects (though not for structs), so this would be nothing new. It's important that we have them, because otherwise common dynamic dispatch would be very annoying to write.
2. Structure layout is already not defined. The compiler is permitted to reorder structure fields as it likes. However, you can force it to adopt your specified in-memory order with the `#[repr(C)]` annotation.
> It would be nice to have a new C, finally. But the more a language has an opinion on runtime layout, behavior, and symbol names, the less C-like it becomes.
The language already has an opinion on runtime layout and symbol names. However, you can specify the layout and symbol names manually if you like (through `#[repr(C)]` in the former case and `#[no_mangle]` in the latter case).
> Rust got this right when it decided that GC was NOT the correct default behavior for a language. The reason you see people playing with OS kernels in rust and not as much in D is, I think, mainly due to this decision. I think rust should continue carrying this torch.
GC has performance costs, while trait objects and symbol names do not, as long as they're opt-in. Garbage collection and virtual dispatch are completely different things; having one in no way moves us closer to the other.
Right, of course not. The comparison was philosophical rather than technical.
My point was that I believe there's a "sweet spot" for a language that is expressive and convenient and modern, but also tries hard not to stray too far from C's spartan abstract machine model (and when it does, it exposes that complexity in a composed pluggable fashion).
I'm beginning to believe rust really has a shot at replacing C and needs to court "bare metal" programmers as well as higher-level programmers to do it; I'm just preemptively registering my wish that rust continue to head down that path.
Prior to rust the hobby OS dev community was primarily C / asm with some honorable mentions for other languages. It's also something of a stand-in for the requirements of the professional embedded community.
For these use cases, it's just really cool to be able to use more and more "layers" of the language as you implement more of the underlying abstract machine model.
[1] http://jvns.ca/blog/2014/03/12/the-rust-os-story/ [2] https://github.com/rust-lang/rust/wiki/Operating-system-deve...
Not exactly.
Any program even in a GC'd language can allocate non-GCd data. Even in Java, there is the Unsafe class that can do manual mallocs/frees. C# integrates it with the language. If you do a big pile of work on the non-GC heap then no GC would be triggered and GC is effectively "zero cost" for this code.
And vtable dispatch has a cost for any code that calls a virtual method. Yes, it's "opt in" but this can be misleading. You pay the cost of the virtual method call every time it's invoked. Static languages like Rust and C++ require the programmer to explicitly state which methods can be virtual to try and control this cost, but that isn't the only way.
For example the JVM is capable of eliminating the virtual method dispatch overhead in almost all cases without requiring the programmer to manually specify which methods are virtual. In fact, the JVM can eliminate the vtable overhead for calls that could be virtual, but in fact at that specific call site are not, and even for calls which are only slightly virtual (e.g. there's only really two destinations). So it's possible that in a JIT compiled program you have way more virtual dispatch in theory, but less than the equivalent C++ program would in practice once the code is warmed up.
So vtable and GC performance are very complex topics which no longer reduce neatly down to our intuitions.
So you're basically repeating what he said -- I don't see how the "Not really" you begin with is justified.
Of course you "pay the cost of the virtual method call every time it's invoked".
And you don't pay it any time it's NOT invoked.
That's the whole idea.
In languages like C++ or Rust you pay the cost every time the method is invoked. With other types of compiler you may not pay the cost even if the method is marked as virtual, even if it's actually used virtually in other parts of the codebase, due to call site specialisation.
http://featherweightmusings.blogspot.com/2014/12/my-thoughts...
It also links to the list of RFCs that have been postponed until post-1.0:
https://github.com/rust-lang/rfcs/issues?q=is%3Aopen+is%3Ais...
For instance, you can use Arc [1] to share memory between threads or Rc [2] to enable thread local GC for a specific variable or RefCell [3] to safely share mutable memory between threads. You can even use raw pointers: it's unsafe (and therefore has to be wrapped in an "unsafe { ... }" block) but possible. Finally, you can call C very easily from Rust [4].
[1] http://doc.rust-lang.org/std/sync/struct.Arc.html
[2] http://doc.rust-lang.org/std/rc/struct.Rc.html
Besides, consider the fact that we've written hundreds of thousands of code for working, non-toy projects in the language, including the Rust compiler, crates.io, Servo, etc.
As a retrofit to C++, that idea wasn't going to work. It would have either been too restrictive or unsafe. It had to be built into the language at a deeper level. That's what Rust does.
[1] http://www.animats.com/papers/languages/cppstrictpointers.ht...
Deleted comment