#![feature(scoped_threads)]
use std::thread;
use std::sync::Mutex;
fn main() {
let a = Mutex::new(0u32);
thread::scope(|s| {
s.spawn(|| {
*a.lock().unwrap() += 1;
});
s.spawn(|| {
*a.lock().unwrap() += 1;
});
});
assert_eq!(*a.lock().unwrap(), 2);
}
A nice glimpse of structured concurrency.[1]: <https://doc.rust-lang.org/std/thread/fn.scope.html>
Edit: removed unnecessary "mut" from declaration of "a".
Scoped threads are such a brilliant thing in the context of lifetimes.
Pre-1.0, we used to have JoinHandle<'a> which let you use the borrow checker to its full potential with multiple threads. This is a big practical hurdle today, where mutices are required even for simple cases which don't suffer data races in practice.
use std::{sync::Mutex, thread, thread::sleep, time::Duration};
fn main() {
static MUTEX: Mutex<String> = Mutex::new(String::new());
let t1 = thread::spawn(|| {
MUTEX.lock().unwrap().push_str("hello ");
});
let t2 = thread::spawn(|| {
sleep(Duration::from_millis(10));
MUTEX.lock().unwrap().push_str("world!");
});
t1.join().unwrap();
t2.join().unwrap();
println!("{}", MUTEX.lock().unwrap());
}
Of course you can't really use it for everything (ie. whatever you put inside `Mutex::new()` needs to be static as well, but still, it's a nice option tooSuper cool not having manual thread joining code.
These are synthetic benchmarks but it's quite significant in them.
From a different tweet:
> It's the total time for 32 threads each doing 10'000 lock+unlocks (on a 64C/128T threadripper). So, the numbers you quoted correspond to a lock+unlock operation going from 8.75ns to 2.45ns, under low contention.
> The numbers can vary a lot in different situations/hardware though.
Not always. Mutexes can be really fast (10-20ns), especially since they often optimistically spin, and Arc in Rust is (often) relatively low cost since you can hand out "free" refs without touching the atomic.
If removing the Arc/Mutex would require allocations the Arc/Mutex could easily be faster.
> Not always
Yeah, that's what "often" means.
> Mutexes can be really fast (10-20ns)
Notably, still worse than 0 ns. Ditto for Arc's refcounting and additional allocation. I'm not saying go on a crusade against Arc+Mutex here, but the easiest way to make effective use of modern multicore CPUS is to go to shared-nothing, independent data-per-thread designs (obviating Arc+Mutex). And if you aren't using Arc+Mutex, it's harder to accidentally share mutable state between threads.
Modern fast mutexes are perfect for that, because their uncontended case is so good. This also inculcates the correct choice for the programmer, you should prefer to write code that is less often contended, not fight hard to get better contended performance at a cost of worse uncontended performance. Contention is bad even if your mutual exclusion primitive performs well.
But Mara measured across simulated workloads with varying contention and this fix improves them all to different extents.
Because it's an incredibly efficient, safe option for doing so. Lots of shared state is rarely contended. For example, imagine you have a 'Config' that gets updated periodically in the background, readers of that config only check for updates every 1 second, and you have 7 parallel readers (and 1 writer for an 8 core system).
A Mutex is a trivial way to solve that problem that will be extremely efficient.
The worst case for an atomic write is two additional cache line flushes, iirc.