Why Rust? [pdf]
oreilly.com
oreilly.com
http://www.oreilly.com/programming/free/why-rust.csp
"One of the original designers of the Subversion version control system, [Jim Blandy is] a committer to the SpiderMonkey JavaScript engine, and has been a maintainer of GNU Emacs, GNU Guile, and GDB."
> But all those other languages include explicit support for null pointers for a good reason: they’re extremely useful. [....] The problem with null pointers is that it’s easy to forget to check for them.
There is no inherent reason for that; just that mainstream languages which support `null` haven't been checking `null` usage. Some of the recent ones do. For example, Kotlin has non-nullable types by default, and null types have to be explicitly marked and checked for null-ness.
Even for Java, null analysis is built into Eclipse with the help of annotations. Though, ofcourse, it would be far nicer to have it baked into the language.
What's important is that you can't use a value of type Option<T> as if it were T; you have to check it first. This is helpful for non-pointer types as well; I include an example of that.
Without collapsing that indirection, Option<Option<T>> is perfectly expressible:
Just (Just foo): <ptr> -> <ptr> -> foo
Just Nothing: <ptr> -> <0>
Nothing: <0>
It's true that some languages expose "nullable references" as a distinct type and not a detail of representation, and it's true (... and sometimes obnoxious) that this doesn't layer as nicely (in particular, it's not a functor).For example, `Option<i32>` is probably (the compiler gets to choose) going to be represented as two four-byte values: the discriminant, which distinguishes the `Some` and `None` cases, and then a space for the value `v`, for when the discriminant says we have `Some(v)`. Since zero is a perfectly fine value for an `i32`, we have to store the discriminant separately.
But note that this is just a flat eight-byte value. There's no heap allocation involved. It's just as if you'd written in C:
struct O { enum { Some, None } discriminant; int32_t value };
I compiled a program that uses `Option<Option<i32>>`, and looked at the DWARF debugging info to see what the compiler did with it. It seems to represent this as a twelve-byte value: four bytes for the discriminant for the outer `Option`, followed an eight-byte `Option<i32>` value laid out as before. Since you can get the address of a value held by an enum, I guess this makes sense; the compiler can't combine the discriminants or do anything clever like that.Option<i32> could have discriminant values 0 or 1, and Option<Option<i32>> could use discriminant value 2 for None, and 0 or 1 mean the object's an Option<i32>.
Edited to add: With your proposed encoding, the memory contents of an Option<Option<i32>> is identical to an Option<i32> precisely when there is an Option<i32> to speak about. And when there is no Option<i32>, you can tell that with one comparison, rather than by backing out an unknown number of levels. I like it.
I don't think so. It's just a pointer to a pointer, rather than a single pointer. That is,
int **value;
and either value == null
or *value == null
or **value == some value of interestSay you have a function, `f : (bool, T) -> Option<T>`.
function f<T>(cond : bool, value : T) : Option<T> {
if cond {
return value // equivalent: Some(value)
} else {
return null // equivalent: None
}
}
let x = f(false, f(false, 1)) // equivalent: None
let y = f(true, f(false, 1)) // equivalent: Some(None)
There is no way to tell apart `x` and `y`.Hm... not sure if I am. Or, at least, `null or T` is not a valid representation for type `Option<T>`. It's all very confusing.
> and your function is not well typed.
How so?
With AnimalMuppet's approach, you would need to return a pointer to a T (so, `&value` in something Cish). Note that this is not the same thing as a nullable reference in, for instance, C#.
`null or T` is a valid representation only when T is a non-nullable reference. It doesn't work when T is optional, and it also doesn't work when T is a primitive type, a struct, &c.
That said, there's obvious reasons it might be wanted an optimization where it's possible. For that specific case, `Just value` could be specialized to `value`, but that can be done by the compiler (akin to automatic unboxing, elsewhere).
The C++ case where a distinction between `foo = null` and `*foo = null` exists is indeed much closer to a an option type. You're right to point it out.
We now need a "programming in Rust" book which is not by one of the Rust developers, who are too close to the design.
I've assisted Jim in giving presentations on Rust at OSCON this year, and trust me when I say that he has no qualms about criticizing the language when he sees fit. His perspective is definitely that of an outsider's, and like you I agree that that's ideal from a pedagogical point of view.
That's not entirely true. There are some mechanisms that you can't really do in C++, at least not without a lot of boilerplate. And Rust's generics aren't quite as powerful (as in metaprogramming) as C++'s yet. I hear they're working on it, though.
It's easy to define a list of integers in pretty much any language,
class IntList { void addInt(){...} }
Most languages let you make a generic List, class List<T> { void add(T){...} }
The thing that's cool about HKT, is it lets you swap the other side, class T<Integer> { T operate(Integer) }
So you can deal with a Set or a List or an Optional or whatever might want to deliver integers. It seems a little goofy, but it's sort of like a Super Lambda. an anonymous function can only really do its one thing, but this lets you bundle a bunch of related lambdas together, so if you have, say Pay, you can group together all the different types of tax that are applied to pay, calculate total tax, etc. And it works over whatever random data structure you wantIt seems a little crazy at first, but with clean syntax, it's a fantastic way to split the work apart from the way the data is represented.
* what kind of animal should it have on the front?
* will it tell me 8 different, confusing, obsfucated ways to do something that would be much faster with printf?
* will it do that in chapter 1?
* ???
* profit?
The std::thread::scoped function used here is undergoing some redesign...
/grumble /grumbleYes, I know, it'll all eventually settle down, but darn its annoying at the moment trying to use stable.
(The 'crossbeam' and 'scoped_threadpool' crates have different implementations of this idea, though.)
While this is probably factually true -- I've had quite the opposite experience. :(
- I want to write benchmarks for my code, but I need `#[feature(test)]` AKA nightly.
- I want to [format my Rust code](https://github.com/nrc/rustfmt), but I need nightly.
- I want to use Huon Wilson's SIMD crate, which requires nightly.
I'm not saying these things should be stable right now, but I am saying if you try to use Rust right now you _will_ encounter a lot of stuff marked as unstable and/or only usable on nightly due to XYZ. I still think Rust is amazing and everyone should give it a shot though, but it's not as stable/polished experience as I think it will be in say a year from now. The language is stable, but the ecosystem just isn't yet.
P.S. You've helped me personally all too much over the Rust IRC channel, so thank you a million times for that! :)
I also think that year-from-now Rust will be way nicer than at the moment, it's absolutely early days. But there's a lot you can do, even on stable, right now. This is why I said it depends on what you're doing. Some use cases have absolutely no choice but to use unstable things. Others have options, and others only use stable.
As to your three points:
1. You can configure things so that you only need nightly when doing benchmarking, and run on stable the rest of the time.
2. Compiling `rustfmt` needs nightly, but using it doesn't, as it's a binary.
3. Yeah, the SIMD stuff just happened, so it's gonna take some time.
Developers like to say they want choice when it comes to formatting, but take it away and you're left with resigned programmers who are very productive at reading code.
^^ Perhaps not the right outlet for this, but maybe you can point me to where this discussion is taking place.
Also, rustfmt is a binary which can work without nightly once you've compiled it.
I built rustfmt with nightly Rust and it dynamically links all kinds of shared objects from the Rust install. This doesn't happen for a simple hello world so there must be something different in rustfmt's build.
Stability mainly matters when you release a library and it stops working on a newer compiler, breaking everyone downstream. That's a pain for the downstream users, because they need not necessarily understand your library internals, and the software they're building will be completely broken until you update it.
If an optional developer tool breaks, it's only you that's affected, not downstream users. You can wait for it to be fixed, no problem.
Rustfmt uses internal Rust APIs to parse and reason about the code. It could use something else (like syntex), but that would be more work.
That doesn't mean that everything that's available on nightly will be available in stable in seven weeks, of course. Everything lands as stable and then is made stable at some point in the future, at least one full cycle after it's landed. There's lots of pre-Rust 1.0 stuff that's still nightly-only.
For example, try using https://github.com/serde-rs/serde without nightly.
Possible? Totally.
A right pain compared to using nightly? That too.
I'm also a huge fan of the Chapel Programming language. Anyone else think they're both hitting a sweet spot and would potentially have a beautiful love child?
It may come as a surprise to some but most systems engineering has very little hard computer science involved and is mostly about achieving very simple tasks in reliable ways with very robust error handling.
You rarely implement novel data-structures or algorithms. Not shipping with cyclic data-structures in Rust is not really a problem for most systems applications. That said, you actually -can- build them in Rust, and it's likely there will be nice libraries with various different allocation/memory management and other runtime tunables and performance tradeoffs.
If you exclusively consider hard computer science problems as "hard problems" then no, Rust is probably not for you. Consider Haskell or Julia.
> What are the technical reasons you can't guarantee an absence of leaks?
Well, 'leak' is one of those things that's easy for a programmer to understand, but a hard thing for a computer to understand, because it's really about intent. How long did you intend for some resource to live? Any global value is, in some sense, a leak. We had a long discussion about this, and, at least currently, we couldn't come up with a formal enough definition of 'leak' to even start tackling the problem of "how do we solve leaks." (It is entirely possible that I am unaware about research on this topic... but given that solving leaks wasn't a goal of Rust, fixing it would just be gravy anyway. You have to choose your battles, and Rust certainly isn't perfect.)As for how safe Rust can leak:
let x = Box::new(5);
std::mem::forget(x);
which can itself just be implemented in safe code: use std::cell::RefCell;
use std::rc::Rc;
fn forget<T>(val: T) {
struct Foo<T>(T, RefCell<Option<Rc<Foo<T>>>>);
let x = Rc::new(Foo(val, RefCell::new(None)));
*x.1.borrow_mut() = Some(x.clone());
}
or something like "I have a thread that is holding the receiving end of a channel that infinitely loops without reading anything off," in which case anything sent down that channel leaks.It is technically possible to guarantee the absence of leaks by introducing a `?Leak` trait to the language. There was a point in time when many wanted Rust to go that way, but it was ultimately decided against (it complicates other things).
this is probably a silly question because the content, though it might always be in RAM, will be branched to an order-of-magnitude less than it'll be used (once it's been loaded to cache)...but still, there should be an observable temporal-boundary to the context stored in cache and branch predictor in a procedural language, that would be manifest as an abstraction-boundary in a more parallelized environment
And frankly: Your definition of "hard" is a bit restrictive.
not a fan of C++ but I personally think having to wade through the building blocks of C was an immensely valuable reverse turing test (ability of human to exhibit human-like behavior? apparently this term is already taken but you get my idea) and I bemoan the future where everyone is protected from their mistakes. How else did one build motivation to do it right? The gamble of a mistake is much more fun than of a compiler error.
Like a fertilizer that stimulates growth while young but a crutch which stunts advancement in old age.
??? These would be super easy to build on top of Send/Sync and lifetime tracking.
You might be confusing Rust the language with a hypothetical Rust/OTP or Rust/HOL or something.
Without those, we're going to see a hundred different solutions crop up, none of them being interoperable.
I'm also sort of hopeful that higher-kinded types can do something about writing code that can run on any such system. If you can have some functions on an abstract Promise<T> type, like fn then(Promise<T>, fn(T) -> Promise<U>) -> Promise<U>, it doesn't super matter which backend you're using. (As I understand it, JS is already here, because they can duck-type their promise spec.)
https://crates.io/search?q=graph
> No support for trying
do I need to have functional experience for exposure to that? Only thing you listed that I haven't heard of.
I would not be shocked if someone figured out a brilliant way to add dependent types to Rust 1.x within the next few years.
You might want to check that the property "the number of black nodes from the root to any leaf is the same" is always valid. With unit testing, you'd implement the insertion and deletion functions, and write a unit test where you give them a few trees and make sure that the insertion and deletion functions don't break that property, probably by looping over every leaf and comparing the black-height. If this passes, you know that your functions are probably correct and at least not obviously broken, but you don't know that there isn't some edge case you haven't thought of.
With a language that supports proofs, you'd tell the compiler about the concept of black-height, and the type of a red-black tree would include the black-height. Therefore, in order to construct a valid red-black tree, you have to have constant black-height. Just as much as you couldn't construct a red node with a red child, you can't construct a node with two children of varying black-height. When you say that your insertion function returns a red-black tree, as part of type-checking, the compiler makes sure that these properties hold -- because in turn, every helper function you call, every constructor, etc. also requires these properties to hold so that they can return an object of type red-black tree.
The downside is that it's much more involved than a few randomly-selected test cases to convince the compiler that you're upholding the red-black tree rules in every single case. This textbook has an example of doing this in the Coq proof assistant (search for the section labeled "Dependently Typed Red-Black Trees"):
QED? Really? I'm not so sure I trust the author to give rust fair treatment anymore. An operating system that does multithreading is not the same as a modern machine.
Edit: Any brave down voters want to explain why? Threads are a way to model concurrency. There are other ways.
If you don't want to come across that way, give us some explanation of what you're thinking and why. Give us some evidence that your statement is true. Something besides just dogmatic assertions that you're right and the author of the article is stupid and/or biased.
Moving on...
In the context of the OS, multithreading (or something essentially equivalent) is the only way to exploit the abilities of "modern" machines. But "modern" doesn't mean very modern (at least, as I understand it). It means the difference between MSDOS and Windows: "Windows is multitasking, whereas DOS... DOS is serially multitasking." (I don't recall who said that, but it was brilliant.) Without this capability, you're only running one program at a time.
Now, I don't care if you implement this as "theads" or something else, but if you don't have it, your OS is pretty much worthless, and has been since about 1990.
> In the context of the OS, multithreading (or something essentially equivalent) is the only way to exploit the abilities of "modern" machines.
I noticed that you qualified multithreading with "or something essentially equivalent". I guess you would agree that qualifying it this way is a good idea. Mr Blandy did not care to do that though: "[...] is the only way to exploit the abilities of modern machines."
I don't like that because it sounds like he's trying to popularize one approach to concurrency without even acknowledging that it is just one approach.
Note as well that Rust intends to have world-class support for a wide variety of approaches to concurrency and parallelism. Threading and channels are already in the stdlib, and fork/join and SIMD are in the works now. The goal is to enable the programmer to be able to choose the best tool for the job.
Edit: No that's not exactly true http://stackoverflow.com/questions/7005759/erlang-on-multico... But they do spawn a processes for each core.
Each thread runs a scheduler and the many Erlang processes are balanced between the schedulers. There is still only one OS process though.
http://erlang.2086793.n4.nabble.com/Some-facts-about-Erlang-...
In any case the distinction between "multiple processes" and "multiple threads" here is fairly minimal, at least on Linux or OS X. As far as the kernel is concerned, they're the same object (Linux calls them threads; Mach calls them tasks). It's just that sometimes, groups of these things share pointers to things like memory map or file descriptor table, and sometimes they have their own pointers. There isn't a whole lot of difference, either at the language level or the kernel level, between running multiple processes from the same executable image that map a shared heap versus just running multiple threads.
I could see Rust's ownership and multithreading amenability being extremely useful at this level, but admittedly I don't know enough.
(And yes, Rust's threading safety applies equally well to that. It doesn't care about the details of how code is running concurrently/in parallel, just that it could be.)
If you're talking about systems programming threads are by and large the main mechanism for concurrency.
Only in really odd architectures do you see things like manually DMA'ing to separate execution units(ala PS3) and are certainly the exception to the rule.
I believe it was Windows that really popularised multithreading as a form of concurrent programming as process creation is so expensive on that platform. I recall Linus ranting against adding threading to the kernel on more than one occasion.
When I went from EXEC 8 to UNIX in 1978, my main observation was that the I/O was better but the CPU management in UNIX was much worse.
Here are Djykstra's P and V functions for the UNIVAC 1108, to run in user space. Written in 1972.[2]
I once added support for threads to Pascal for that machine. Had to add per-thread stacks and manage non-contiguous stack growth.
[1] https://ia601603.us.archive.org/27/items/bitsavers_univac110...
You can see Aaron's initial work on the basics (memory management) on his blog[1], and a higher-level introduction into the power of Rust's concurrency on the main Rust blog (also written by Aaron)[2].
[1]: http://aturon.github.io/blog/2015/08/27/epoch/
[2]: http://blog.rust-lang.org/2015/04/10/Fearless-Concurrency.ht...
It is possible. For example, you could just remove threads and synchronization primitives. :)
But it's not easy -- IMO, compared to data races it's much harder to eliminate deadlocks statically without restricting expressiveness too much. It's even harder if you count I/O related deadlocks as deadlocks (e.g. deadlocking reads on Unix pipes).
In general, any of these properties can be checked statically, just like czwarich said. You have to be conservative, but that's not any different than a type system, or Rust's borrow checker. There are certainly languages that enforce termination, and you can design systems that enforce higher-level progress properties (such as absence of deadlock).
The major difference between a liveness property and the other kind (called a safety property: at no point does this bad thing happen) is that you can't check for liveness properties dynamically.
The real goal is to find a strategy that works well enough, often enough. Dynamic GC gives up perfect collection (aiming to be merely good enough). Rust's data race guarantees prevent safely expressing some subset of correct programs.
Whether deadlock prevention can be done nicely enough for general case usage is an open question.
"Check out our awesome new car! It makes driving 'easy and safe'!"
"Does it prevent crashes?"
"No, that's impossible! But here, look at our onboard computer that prevents many types of driving errors."
I like Rust. I want it to succeed. But I think the rhetoric gets ahead of the language sometimes, and not prioritizing higher level concurrency tools because you have the borrow checker is a mistake.
For example, here's a post of Niko Matsakis' from Feb 2014 about changing a fundamental aspect of the borrow checker in order to better support data parallelism in the future: http://smallcultfollowing.com/babysteps/blog/2014/02/25/rust...
For another example, here's an RFC from Nov 2014 about fundamentally altering Rust's concurrency support in order to better support fork-join parallelism in the future: https://github.com/rust-lang/rfcs/blob/master/text/0458-send...
I'm also confused about your perception that the rhetoric gets ahead of the language. The type system does indeed make things easier and safer. When people ask what that means, we're eager to elaborate on the precise guarantees that Rust provides. Misleading people as to Rust's capabilities is not on the agenda.
I think it's a fair characterization given 1.0 shipped without them, and there doesn't seem to any timeline for standardization (please correct me if I'm wrong).
I don't think I have to tell you that concurrency is one of the most important challenges in modern programming. But all Rust gives you today (and for the foreseeable future) is a pthreads wrapper.
>I'm also confused about your perception that the rhetoric gets ahead of the language
The parent comment says verbatim "Rust goes out of its way to make it easy and safe to write multithreaded code". This apparently doesn't include preventing deadlocks, a problem that is common, hard-to-avoid, and difficult to recover from. Does that not strike you as a bit of an overreach?
Deadlocks are by far the easiest concurrency problem to debug, since it's very obvious when your program is deadlocked, and in most cases a stack trace of the involved threads is sufficient to debug and fix a deadlock. Also, if your program is deadlocked it won't corrupt user data.
Rust's std::sync::Mutex is fortunately non-recursive, which makes it easier to find deadlocks during testing.
Data races are far worse since they may cause arbitrary effects at a later point in program execution, so they take a lot of developer time to track down and very likely lead to data corruption.
From the example, from a deadlock the locked threads and their stacktraces can be determined, which will help in pointing towards where the cause of the problem is. Compared to situations where the problem occurs due to a data-race causing an unexpected/invalid state but a problem doesn't manifest until later. Fixing this type of problem is more problematic as usually any exception/stacktrace that might crop up does not relate any information as to how the state became invalid in the first place. I tend to see these types of bugs more than deadlocks and in my experience they're always more involved in debugging compared to deadlocks.
> But all Rust gives you today (and for the foreseeable future) is a pthreads
> wrapper.
... Plus significant static guarantees about the safety of using said wrapper.But overselling it ("Rust goes out of its way to make it easy and safe to write multithreaded code" except oh yeah it can deadlock at any time) is a mistake.
Writing nontrivial multithreaded code using thread and mutex primitives is very hard to get right (not only due to data races, but also race conditions and deadlocks). Has your experience been different?
>Show me a language that statically prevents deadlocks
OK, idiomatic Go and Erlang will never deadlock. Sure, you can use mutexes in Go, but unlike in Rust they aren't the only means of achieving high levels of concurrency (in fact, they are explicitly discouraged).
I find this attitude from the Rust community disheartening. Not everything needs to be statically verified to be useful.
https://talks.golang.org/2012/concurrency.slide#46
This is great for a couple reasons:
It's automatically concurrent without explicit pooling and locking.
The code flow remains sequential. No callback hell!
It reduces concurrency errors. The locking pattern is naturally acyclic! Nothing races!
Ah, but Go has a big runtime, you say! We can't have that in Rust! Well, here's the same thing in C and in C++:
http://libmill.org/tutorial.html
http://www.boost.org/doc/libs/1_59_0/libs/coroutine2/doc/htm...
Coroutines are a wonderful tool for building concurrent applications, and I dearly hope we get them in Rust.
Forget to unlock a mutex when dealing with shared data? Deadlock (and Go doesn't have RAII mutexes, so it's very easy to do this).
The pattern you put forth is possible in Rust too. Use mio if you want goroutine-like efficiency. Nothing new.
To be fair, how long would that take to debug?
>Forget to unlock a mutex when dealing with shared data?
You'll note there are no mutexes. RAII mutexes are great (although less useful without exceptions). But the entire point of Go's concurrency model ("share by communicating") is that you don't need need to deal with them.
>Use mio if you want goroutine-like efficiency
For a personal project, I would. But will any commercial entity use an unstable, unportable library with a single maintainer for critical functionality in their app? Because that's the concurrent IO situation in Rust, now and for the foreseeable future.
> But the entire point of Go's concurrency model ("share by communicating") is that you don't need need to deal with them.
Exactly the same thing works in Rust, and works "better": the lack of other sharing (except by message passing) is enforced at compile time. Rust ensures that other options are available with as much help for correctness as possible.
> But will any commercial entity use an unstable, unportable library with a single maintainer for critical functionality in their app?
Concurrent IO is inherently unportable, and mio has support for the major platforms (OSX, Linux and Windows, with tests run on all) so I don't know what you mean by that. mio won't be the first or last lib with a single maintainer that a commercial entity uses.
(Unstability is of course a perfectly reasonable criticism, and I'm sure it'll disappear as the library ages.)
Write the version with explicit locks and condition variables and we'll see if it's as easy to debug. :P
>Exactly the same thing works in Rust, and works "better": the lack of other sharing (except by message passing) is enforced at compile time.
Which is one of the reasons why Rust can (should) eat Go's lunch. All it's missing are lightweight coroutines and concurrent IO.
And I didn't know Mio supports Windows now — that's good news! If Mio stabilizes and Rust gets lightweight coroutines (probably requiring compiler support), Rust could be the best of all worlds.
In any case, you seem to be ignoring what everyone is saying: Rust doesn't guarantee deadlock freedom, but it still tries to help. Mutexes can be an important building for some things, but they're not the final story. There's atomics and channels in the standard library right now, and now that 1.0 is released, there'll be a growth of even better abstractions.
One of the people working full time on Rust has a PhD in concurrent programming, and has it as a personal goal to make Rust great at it. You can see his initial work on his blog[1], and read his thesis which introduces "reagents" (something he has expressed interest in implementing in Rust)[2].
1.0 is the start for Rust, not the end.
https://en.wikipedia.org/wiki/Deadlock
Either way, redefining the word does nothing to make mutexes and native threads safer or easier to use in practice.
I can't emphasize this enough — data races are not the only kind of concurrent programming error, yet they are the only kind prevented by borrow checking. Go, Erlang, and Node have high level approaches to concurrency that reduce other categories of errors (in addition to providing massive performance benefits vs. naive native threading).
I think the time to add good concurrency abstractions to a language is before 1.0, but I'm in strong disagreement with the community there. And judging by this thread, concurrent IO isn't even on the core team's roadmap. This is not a good sign for a language that sells itself on concurrency!
Look at the situation 10 months ago. How much has it improved?
https://www.reddit.com/r/rust/comments/2l0a4b/do_rust_web_se...
Channels are a step in the right direction, but are pretty limited without coroutines (channels without coroutines are just synchronized queues). Atomics are cool, but they address a totally different problem. Reagents, well, let's see an implementation.
C++ has needed a successor for many years now, and Rust is the best candidate. But lofty claims notwithstanding, the Rust concurrency situation is pretty dreadful.
And focusing on hangs in concurrent programs that are caused by using a data structure called "mutex" doesn't stop one getting exactly the same symptoms via your apparently-perfect channels.
One can use channels to get all four of the conditions required for a deadlock, especially Go's synchronous-by-default channels.
Any time you have a protocol of multiple tasks communicating with each other in some structured way, it's possible to break that protocol and hence have tasks sitting around waiting for messages that aren't coming. Especially in languages like Go/JavaScript/... which aren't powerful enough to model things like session types in their type system, e.g.
https://www.reddit.com/r/rust/comments/3jhd56/session_types_...
> I can't emphasize this enough — data races are not the only kind of concurrent programming error, yet they are the only kind prevented by borrow checking
Yes, that's exactly why the whole Rust community tries to be very careful about using "data race" when talking about concurrency in Rust.
However, it is the case that data races are the worst sort of concurrency bug: they are undefined behaviour and so can lead to arbitrary memory corruption, possibly only appearing a long way from the actual place with UB. Deadlocks and other problems are, by default, much more controlled in their failure modes.
Rust focuses on truly outlawing large classes of horrible problems: dangling pointers, iterator invalidation, data races (and all without requiring a garbage collector, although a GC barely helps with the latter two). It also tries to help with other problems with fewer guarantees, but even just being memory safe is a huge step up from widely used low-level languages (i.e. C/C++).
All the languages you mention are quite opinionated in their concurrency, imposing costs that Rust doesn't and can't, for its target space. (And, Go certainly doesn't provide any guarantees at all, not even data race freedom.)
> in addition to providing massive performance benefits vs. naive native threading
Important qualification: for IO bound tasks. Which is perfectly fine, but it needs to be understood.
> I think the time to add good concurrency abstractions to a language is before 1.0, but I'm in strong disagreement with the community there
What's so important about being pre-1.0? You seem focused on it, but I don't understand why. What benefit does Rust gain by delaying the release of 1.0 for months/years just to get good async IO support? Why is this particular pet feature any more important than everyone else's pet feature? (There have been so many requests: "why couldn't X make it into 1.0?")
If you're concerned about theoretical fragmentation of the ecosystem... that's not a problem in practice: mio is the standard.
> And judging by this thread, concurrent IO isn't even on the core team's roadmap
Wrong, it's very much on the roadmap, e.g. Alex Crichton (core team member) has been adding windows support to mio himself.
> Look at the situation 10 months ago. How much has it improved?
A lot. There's a burgeoning ecosystem built around mio.
---
In any case, Rust has been stable for barely 3 months. Be patient, and give it time for the concurrency story to blossom from the seeds that have been sown so far. Based on the experience so far, I'm pretty confident that Rust can easily be much better (i.e. more performant and reliable) than both Node and Go and even Erlang. (Of course, it may be syntactically less nice, since those languages bake it in deeply, while Rust is less opinionated.)
I've never used Go, but it has some interesting concurrency ideas. So do Node and Erlang. I naively expect Rust to adopt the best ideas from each.
>One can use channels to get all four of the conditions required for a deadlock, especially Go's synchronous-by-default channels.
One can, yes. But it's pretty easy to avoid cyclic locking patterns when each request is handled by a separate goroutine, as is idiomatic. One thread per request in Rust will drag pretty quickly.
Yes, the borrow checker is an impressive achievement. But is it enough for Rust to succeed? Marketing yourself as safer C++ is what Java already did (with tremendous success) 20 years ago. And the market for systems languages has only shrunk since then (my phone runs Java).
>All the languages you mention are quite opinionated in their concurrency, imposing costs that Rust doesn't and can't, for its target space.
And yet Rust has already partially standardized channels. Finish the channels, add coroutines and you've implemented Go! (I'll note there are already coroutine implementations for C and C++, which do not limit their use as systems languages).
>(And, Go certainly doesn't provide any guarantees at all, not even data race freedom.)
Which, interestingly, hasn't hindered its ability to become a successful language! A lesson worth remembering.
>Wrong, it's very much on the roadmap, e.g. Alex Crichton (core team member) has been adding windows support to mio himself.
I mean... https://github.com/carllerche/mio/graphs/contributors?from=2...
Could be worse, could be better?
>Why is this particular pet feature any more important than everyone else's pet feature?
I don't think "good concurrency is a pet feature" is the winning argument here.
>I'm pretty confident that Rust can easily be much better (i.e. more performant and reliable) than both Node and Go and even Erlang.
Me too! But I'm not sure that's going to happen with a single core developer on Mio and no timeline for standardization.
> I've never used Go, but it has some interesting
> concurrency ideas. So do Node and Erlang. I naively
> expect Rust to adopt the best ideas from each.
This is where your naivete shows. Rust originally did have the same thread model as Go baked into the language and standard library, and it labored for years to find a usable compromise between Go's green thread model and the native threading model. And a compromise is indeed necessary, firstly because we don't just need another Go, and secondly because Go's threading model imposes horrific costs when trying to interoperate with non-Go code (literally thousands of times the overhead that you'd expect). For a language like Rust that intends to interoperate with the native ecosystem, that overhead is unacceptable. After about three or four complete redesigns and rewrites the entire green threading infrastructure was chucked to the curb. Fortunately, Rust is low-level enough that libraries like mio can pick up the slack on their own, and in the meantime libraries that don't need green threads don't have to pay the price.Ehhhhhh not quite. Go provides one threading API, and it's green threading. Rust tried to provide the both green threading and native threading using identical APIs. That was a unique and in retrospect quixotic decision. Most of the problems identified in the RFC stem from the unified API issue:
https://github.com/rust-lang/rfcs/blob/0806be4f282144cfcd55b...
I also understand Rust's implementation was backed by libuv, which, being designed for Node, was a poor fit for Rust:
https://plus.google.com/+nialldouglas/posts/AXFJRSM8u2t
You're the third person in this thread to tell me that Go-like concurrency requires a big runtime, and it remains false. Here are analogous concurrency implementations in C and C++:
http://libmill.org/tutorial.html
http://www.boost.org/doc/libs/1_59_0/libs/coroutine2/doc/htm...
And because of that fallacy, the future of concurrency in Rust is a one-man show, third party library. It's a tremendous loss.
(Concurrency in C and C++ is also third-party libraries... The whole point of languages like Rust and C++ is that powerful functionality like this can be built externally, so that different trade-offs can be made. Languages like Go and Node force one approach, and so when you need something outside it, you're forced to do something suboptimal.)
Without the documentation, stability, portability, quality guarantees, and compiler support (that's a big one — code generation for coroutines needs to be good) of a standard library.
>Concurrency in C and C++ is also third-party libraries
C++ is on track to standardize concurrent file and network IO. Draft specifications have already been published, and Microsoft shipped coroutines in VS 2015. It would be a damn shame if C++ got concurrent IO before Rust.
I would like to use Rust professionally, and I'm sure you do/would as well. But no one can possibly sell their boss on using a project with a single part time maintainer to provide critical functionality.
That said, I do think a much better style of coroutines for Rust would be a C#-esque async/await transformation, converting stack-frames/local variables into an enum, allowing literally zero-cost coroutines (all the state is stored inline, no need to allocate a separate stack). Relevant issues:
- https://github.com/rust-lang/rfcs/issues/388
- https://github.com/rust-lang/rfcs/issues/1081
I'm pretty sure this is quite non-trivial to implement automatically.
---
C++ has had 20 years of stability, Rust only 3 months. Rust will get concurrent IO before C++ has on that time scale.
The goal with 1.0 was to stabilise enough of the language that people can start using it to write libraries that work into the forseeable future, allowing them to seriously explore the space of, for example, concurrent IO in Rust. Once enough exploration has been done (maybe you think enough has been done for async IO now), the functionality can start to become more official.
You can open VS 2015 today and use C++ coroutines backed by Microsoft (and their compiler, which is developed alongside their standard library).
And I am by no means saying that Goroutines are the final story in concurrent IO. Stackless coroutines in Rust would be a dream.
>C++ has had 20 years of stability, Rust only 3 months. Rust will get concurrent IO before C++ has on that time scale.
Concurrent IO is a hell of lot more important than it was in the 1990s, and the relative timescale is irrelevant for people choosing between Rust and C++ today (or Go, Scala, Clojure, C#, etc.).
>Once enough exploration has been done (maybe you think enough has been done for async IO now),
Exactly the opposite — I think the number of developers working on this (the Mio author plus Alex Crichton, maybe some offshoots) is far too few.
And the attitude I'm seeing from some core developers in this thread (concurrent IO is a "pet" feature that the community will someday deliver fully formed and ready for "blessing") is a huge disappointment.
There are cross-cutting concerns that apply to everything (including concurrent IO libraries) that development work is focusing on. Rust doesn't need to eat Go's/Scala's/Clojure's... lunch 3 months after it was released, taking a year or two to settle in and branch out seems fine to me.
Well, what domains is Rust targeting? Rust still isn't a good fit for embedded/kernel stuff without allocators and OOM handling (not to mention every architecture not targeted by LLVM).
That leaves userland applications, and how many applications don't need concurrent IO?
I'm definitely being impatient, and I'm sorry for that. It's just frustrating to see only one core member working on it.
Many user-space applications don't do heavy network work, which is where concurrent IO is most necessary. E.g. games or a scientific simulation, even web-browsers don't need of concurrent IO (they're not trying to juggle thousands of connections).
Which doesn't, correct me if I'm wrong, provide a stable allocator API or OOM handling yet.
>E.g. games
When your client pings 400 servers, how is it doing that? How is the game server implemented?
>a scientific simulation
That runs on one machine, and doesn't do a lot of disk IO? (concurrent file IO matters too).
>even web-browsers
Open the Chrome dev tools "network" tab, and then open Gmail. How many requests did it make?
Again, concurrent IO is important.
> I think it's a fair characterization given 1.0 shipped
> without them, and there doesn't seem to any timeline for
> standardization (please correct me if I'm wrong).
The release of 1.0 was not an indication that the language was 100% complete, or that the stdlib was 100% comprehensive. The 1.0 release represented a stable foundation upon which to build an ecosystem. Going forward, the Rust developers absolutely do care about providing more concurrency primitives. Here's Aaron Turon's recent work on implementing epoch-based memory reclamation for implementing lock-free data structures: http://aturon.github.io/blog/2015/08/27/epoch/ . Of the two interns the Rust project was granted this summer, one of them spent the entirety of their time working on supporting SIMD in the language (http://huonw.github.io/blog/2015/08/simd-in-rust/) while the other spent their time specifying the behaviors that are allowed inside Rust's `unsafe` blocks, which includes taking a good long stare at Rust's memory model and all its concurrency primitives and making sure that they're sound. While we're on the topic of interns, you're also overlooking the `Arc` pointer in the stdlib, which is an intern project from more than three years ago which allows one to safely share memory between threads. Meanwhile, pcwalton (Rust core team member and full-time Servo developer) has been working on shmem support and multiprocess allocators for Servo. In addition, the Rust stdlib contained an implementation of fork/join prior to 1.0, but at the last minute the API was found unsound for a few edge cases and so it was deprecated and a working implementation was deferred until later (today there are at least two working reimplementations of this on crates.io with the unsoundness fixed). > This apparently doesn't include preventing deadlocks, a
> problem that is common, hard-to-avoid, and difficult to
> recover from. Does that not strike you as a bit of an
> overreach?
Not at all. Statically preventing data races is an enormous leap forward in the state of the art. Here's a quote from Matthias Felleisen of Northeastern University on teaching Rust to students without experience in concurrent programming: https://www.youtube.com/watch?v=JBmIQIZPaHY&feature=youtu.be..."I had them program parallel programs in Rust, and as some of you know, Rust prevents race conditions with its type checking. [Note that he's incorrect here, Rust prevents data races, not race conditions in general.] I will admit, the first two weeks, because we used the beta release, was a mess. We couldn't understand the type error messages. But once we got over the hump, I was blown away, that these kids never had problems writing parallel programs in imperative style. There were no race conditions. The type system slapped their fingers. We taught them how to design, the type system enforced it, and lo and behold, I hate to admit this but I have to admit it, they just didn't have problems with this stuff."
>Statically preventing data races is an enormous leap forward in the state of the art
The borrow checker is a great achievement. But it doesn't prevent deadlocks. And because Rust doesn't provide higher level concurrency tools than threads and mutexes, deadlocks are going to be a significant problem in practice until such tools are standardized.
Could you describe how you've developed the misconception that Rust does prevent deadlocks? Whatever it is, we should try to fix it right away, because Rust has never tried to prevent deadlocks, and suggesting that it does is not good.
> And because Rust doesn't provide higher level concurrency tools than threads and mutexes
Rust's standard library has channels.
https://news.ycombinator.com/item?id=10192042
Reasonable people can disagree about what "safe" and "easy" mean, but I don't think Rust's concurrency primitives are either given the (very real, very bad) possibility of deadlock.
>Rust's standard library has channels.
Which are a good start, but only a partial solution without coroutines (and the other additions you mentioned).
OK. Then I don't know what to say. An infinite loop is a deadlock. Unless you have a programming language that can guarantee termination, you can't prevent deadlocks statically.
So yes, I think a brief marketing pitch of "safe and easy concurrency" is totally appropriate given the domain Rust is shooting for. In fact, it is one of Rust's strengths, so to not advertise it as such would be quite odd.
> Which are a good start, but only a partial solution without coroutines (and the other additions you mentioned).
OK. But you said "And because Rust doesn't provide higher level concurrency tools than threads and mutexes" which just isn't true. I'm trying to clear up what it is Rust does have. I'm not saying it has everything. Writing software takes time.
It isn't. I don't know what else to say:
> not prioritizing higher level concurrency tools because you have the
> borrow checker is a mistake.
Ahh, but to mix up your analogy here, the borrow checker is the engine, and the higher level tools are like building fancier cars. You have to get your foundations built before you can build higher-level things on top of them.Now that we have a foundation, these things are starting to appear. See https://news.ycombinator.com/item?id=10131429, for example.
Specifically, it's great that Rust has a capable atomics library. I'm sure we'll see many good concurrent data structures come from it. But that's not quite what I meant by concurrency tools. The thread primitives Rust offers are the same ones POSIX standardized twenty years ago. It would be very valuable to have something like OpenMP (parallel loops, etc.) and something for concurrent IO (either asynchronous or coroutine-based).
There's no reason Rust can't have the performance of C++ and the easy concurrency tools of Go or Erlang.
(And the language can't have exactly the same stuff as Go or Erlang without adding a runtime, which is contrary to the goals of the project.)
>(And the language can't have exactly the same stuff as Go or Erlang without adding a runtime, which is contrary to the goals of the project.)
Here is 90% of what Go gives you as a C library. No heavyweight runtime necessary:
Rust has channels in its standard library. Admittedly, they do not have the same functionality as Go's concurrency primitives. The two most significant omissions are probably a multi-producer/multi-consumer channel and a more complete (and stable) `select` construct.
There is also my `chan` library, which replicates Go functionality with respect to channels: http://burntsushi.net/rustdoc/chan/ --- Notably, you still have to use native threads.
Your `libmill` example is a coroutine library with a scheduler and everything. That is definitely too much runtime for standard Rust. However, there are people working on coroutines in Rust, on which something like `libmill` could be built: https://crates.io/search?q=coroutine
Is it incompatible with native threads? (no) Does it affect non-coroutine functions? (no) Does it have any effect whatsoever when not using the library? (no)
I encourage you to look over the code before dismissing it as "definitely too much runtime."
I encourage you to take a look at ongoing work on coroutines in Rust, which was the part of my comment where I didn't dismiss your point.
I think the portion of my comment you quoted was taken out of context.
https://crates.io/crates/mio for the async IO.
There was another library out there that provided tons of concurrency utilities, but I can't find it now. Parallel loop-like syntax should be easy with scoped_threadpool and a macro though.
Rust tries to avoid stuffing everything into the standard library. So the lack of utilities in the stdlib isn't an artifact of some ignorance of concurrency, it's because Rust doesn't want to keep everything in the stdlib. Rust keeps the basic framework for thread safety in the stdlib (though it doesn't need to, not exactly), but the rest is built upon by the community.
Ruby has EventMachine, the standard lib, and Celluloid. None of these are interoperable.
Python has the standard lib, Twisted, and a number of other projects. All have the same problems.
C++ has Boost, Asio, a few others. This is more along the lines of where Rust is headed.
C has /countless/ options. None of them are standard, and it's totally understood in such an old language.
---
Go has goroutines: it's a harmonious, unified ecosystem. I greatly dislike go, but this is one thing they do absolutely correctly.
Erlang has fully-preemptive multitasking (which the entire ecosystem is built upon).
Haskell has incredible parallelization and concurrency primitives built right into the stdlib.
Certain things belong in the stdlib. Async IO is definitely one of them. If not an implementation, a defined, common interface of some nature, to help prevent some of the fragmentation.
async/await, if I am not mistaken, will require compiler and borrow-checker support. This would be a good start.
That's a valid philosophy, but also one that leads to problems with fragmentation, quality, portability, dependency management, and compiler support.
Concurrent IO is surely important enough to standardize.
No, not necessarily. We would like for the standard library to remain minimal, but one of its important functions is to collect common interfaces to maximize interoperability between crates. This has worked well in practice so far.
Similarly, the standard library provides portable facades on top of platform specific APIs, for example, for performing IO. Crates can take advantage of this so that they can be portable themselves. Moreover, crates themselves can also provide portable facades over platform specific APIs, so I'm not convinced that this will be a problem in practice.
Dependency management is handled quite well by Cargo. It has been a wonderful tool to have at our disposal and is really the crux of what makes a small standard library possible.
Compiler support is a good argument, but one that I hope becomes weaker in time as we stabilize more functionality.
To be clear, I agree that a small standard library has its own downsides. In particular, quality is IMO on of the best arguments against a small standard library. Fortunately, we trying to mitigate this by adopting officially blessed libraries into the `rust-lang` organization: https://github.com/rust-lang/rfcs/blob/master/text/1242-rust... --- This allows us to avoid the problems with a big standard library (too much stuff that is hard to evolve because of stability) while still providing quality with crates that we promise to maintain.
Historically, it's been a massive problem. See my other post. Rust is already seeing IO-related fragmentation.
C, C++, Ruby, Python, etc. are all massively-fragmented ecosystems regarding IO, concurrency, (safe) parallelism. I don't think we need to soil a great language (Rust) with these same mistakes.
Also, none of those languages started out with a tool like Cargo.
I made a few other comments about mitigating this as well that should be considered.
I'm not suggesting a Python-esque stdlib. I'm suggesting a multi-threaded, cross-platform, (ideally, edge-triggered) event system, above which higher level primitives can be introduced and safely interoperate.
If Rust already has std/net, this is not that far of a gap to close. Granted, implementing a reactor or (higher level) green threads greatly affects the way your programs execute, in my opinion, the benefits of a "blessed way" would outweigh the problems.
Also, as with all of those languages with fragmented ecosystems: nothing is preventing a developer from implementing their own solutions, they're just heavily encouraged to be compatible.
> To be clear, I agree that a small standard library has its own downsides. In particular, quality is IMO on of the best arguments against a small standard library. Fortunately, we trying to mitigate this by adopting officially blessed libraries into the `rust-lang` organization: https://github.com/rust-lang/rfcs/blob/master/text/1242-rust... --- This allows us to avoid the problems with a big standard library (too much stuff that is hard to evolve because of stability) while still providing quality with crates that we promise to maintain.
I could absolutely see one of these crates providing async IO. I am less sure of seeing it wind up in std. Some examples of crates that are currently on track to being blessed---but maybe never end up in std---are regex and rand.
I disagree with adding good package tooling after-the-fact and using its failure in preventing fragmentation as a reason for why Cargo will be ineffective. Overcoming inertia is hard. Having Cargo at the outset is a nice advantage we have working in our favor. We should acknowledge that.
> If Rust already has std/net, this is not that far of a gap to close.
It's a pretty big gap IMO, especially if you want to provide a common high level interface. std::net is a portable interface around platform specific APIs and not much else. Async IO is quite a bit more involved.
> I'm not suggesting a Python-esque stdlib. I'm suggesting a multi-threaded, cross-platform, (ideally, edge-triggered) event system, above which higher level primitives can be introduced and safely interoperate.
Note that my comment was specifically about arguing against a Python-esque stdlib, or rather, in favor of a small standard library. It was not meant to target omission of any one particular feature, which is what you seem to be focused on. (The criticism I responded to was not specific to async IO.)
I'm looking at the standard library and I see only threads and channels. Is there anything higher level? Parallel map, reduce, etc., something like OpenMP?
That said, I wish they would have looked a little further outside Reddit's homepage rendering for parallelization inspiration. Reddit does not take a long time to render. However, CNN.com [which they looked at too] does, so...