How Rust Achieves Thread Safety
manishearth.github.io
manishearth.github.io
When you pass data to another thread, there are three options: 1) pass a copy, 2) hand off ownership to the other thread, and 3) transfer ownership to a mutex object, then borrow it from the mutex object as needed. All of these are memory and race condition safe due to compile time checking.
Those are the concepts. The Rust syntax needed to support it is somewhat complicated, but if you get it wrong, you get compile time error messages.
This may be the biggest advance in software concurrency management since Dijkstra's P and V. Almost everything in wide use is either P and V under a different name, or some subset of P and V functionality. Locks are not a basic part of most languages; the language doesn't know what a lock is locking. The Ada rendezvous and Java synchronized objects are exceptions. Those were good ideas, but too restrictive. Finally, we're past that.
Go could have worked this way. Go originally claimed to be concurrency safe, but it's not. You can pass a reference across a channel, and now you're sharing an unlocked data object. This is easy to do by accident, because slices are references. Because Go is garbage collected, it's almost memory safe (there's a race condition around slice descriptors that can be exploited), but it doesn't protect the program's data against shared access. In Rust, when you pass an non-copyable object across a channel, the sender gives up the right to use it, and the compiler enforces that.
[1] http://blog.rust-lang.org/2015/04/10/Fearless-Concurrency.ht...
> C++11 standard officially bans data races as undefined behavior.
Relaxed access of atomics is not UB in C++, that would be ridiculous. Clearly one of these two sentences is wrong or imprecise; I would bet it's the first one.
Also https://news.ycombinator.com/item?id=9796245
I think your initial suspicion is right. The next section: "There are several possible definitions of a data race. Probably the most intuitive definition is that it occurs when two ordinary accesses to a scalar, at least one of which is a write, are performed simultaneously by different threads. Our definition is actually quite close to this, but varies in two ways:" ... "Instead of restricting simultaneous execution, we ask that conflicting accesses by different threads be ordered by happens-before. This is equivalent in simpler cases, but the definition based on simultaneous execution would be inappropriate for weakly ordered atomics." continues at: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2007/n248...
This necessitates some level of synchronization.
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2007/n233...
There is a lot more explanation here that should cast some light on the situation: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2007/n248...
A race condition or race hazard is the behavior of an
electronic, software or other system where the output is
dependent on the sequence or timing of other
uncontrollable events.
Example: https://gist.github.com/anonymous/eb72f1091bd1592df552Output:
$ for i in {0..10000}; do ./race; done | sort | uniq
0
1http://blog.regehr.org/archives/490 has a fairly good discussion of the generally accepted definition of data race.
See http://blog.regehr.org/archives/490 for more. But the TL;DR is:
A data race happens when there are two memory accesses in a program where both:
* target the same location
* are performed concurrently by two threads
* are not reads
* are not synchronization operations
I would argue you have a race condition but not a data race.> This may be the biggest advance in software concurrency management since Dijkstra's P and V.
+1 . I'm not sure whose idea it was, but the idea behind the Send/Sync traits is brilliant. It's easy to say "we can mark a type as thread safe and non threadsafe". It takes some thought to come up with a design which works in a world of data sharing. Someone realized that thread safety was actually two, intertwined and interdependent concepts -- thread safe and share-safe[^1], and together they allowed for a good degree of safety without being too restrictive (unlike just having "non threadsafe" types).
[^1]: I am greatly oversimplifying the situation by calling them these names, but .. for a less simplified characterization, read the post :P
There's more than those three: data can be shared via reference or reference counted pointers[Arc], which allows immutable access (more generally, concurrent access to any types for which this is safe, e.g. atomics and mutexes). A mutex is only necessary for mutating (nearly) arbitrary shared data. This is particularly powerful for literally zero-overhead read-only shared memory while still maintaining Rust's safety guarantees.