There's No Such Thing as “Implicitly Atomic”
belkadan.com
belkadan.com
Caches don't break atomicity.
Hint: they're not, and that causes all kinds of subtle problem when programmers assume they are, especially when it comes to atomicity. The compiler caches a value in a register and two threads synchronize on that value? Boom. Or maybe no boom when one was expected.
The example cited was in a library written in C.
This is why it is important to tell the language what you mean. Even in C non-atomic access don't have these guarantees. It may seem pointless because in 99% of cases that "non-atomic" store compiles to the exact same instructions as a relaxed store. But that is just because you are getting lucky. The language doesn't guarantee that and with the atomic store the compiler is well within its right to emit something different (like two stores, spilling dirty values, ...).
> Otherwise, each read of a single-word-sized or sub-word-sized memory location must observe a value actually written to that location (perhaps by a concurrent executing goroutine) and not yet overwritten.
However, it allows the implementation to immediately exit and report an error as well.
So for both Rust and Go you'll get atomic accesses when you need them, and for C etc you only get atomic accesses if you ask for them. Which pretty much speaks to the different programming models: Rust and Go will only compile the subset of possibly-valid programs that can be expressed by the language and proven by the compiler. C compilers will only reject code they can prove is incorrect.
But this gets to the deeper problem of the blog. There are three ways you can reason about these kinds of problems:
1. What happens with the execution of assembly language on the chip? Behavior is pretty well specified by the ISA, and you absolutely can reason operationally. The x86 memory model includes TSO, so you get fairly strong guarantees; it's basically acquire/release "for free," so you know you won't get tearing or other such things.
2. The formal memory model provided by the language. In the case of C and C++, it's also pretty well specified, and there's lots of work that's gone into it over many years. Again, it's possible to reason productively in this world, though it's quite different than case 1 - ultimately you've got causality graphs and other things expressed on an abstract machine. Ideally you'd do this work with formal proofs, but model checkers like CDSChecker can help a lot.
3. Informal reasoning based on an intuitive model of what the computer "should" do. This is serious YOLO territory, and basically a guarantee that whatever you write will be broken, possibly leading to serious security vulnerabilities. It's popular among a subset of confidently wrong HN commenters though, who I expect will come out in force in this thread.
[1]: https://youtu.be/g9Rgu6YEuqY?si=ZvKlDnqKOqzfFVSZ&t=4267
I think it was something about how L1 wouldn't flush down to the inclusive L2 until commanded to, so external reads hit the stale L2.
Alpha being a pain here would be on brand, but I think this was a different arch.
Maybe I'm tired and lack imagination, but please enlighten me. How can splitting up a store speed something up or reduce the code size? Assuming it is aligned and on a 32bit native CPU.
Final 32-bit store effectively stores two 16-bit values or'ed together as high-16 and low-16. You can get this without any shifting and simple addition and the proper values. If these two sixteen bit values are calculated along two very different paths which have different lengths, the compiler might opt to spill one of them early, only to reload the spilled value just before the final addition.
From there it isn't too hard to see some liveness analysis determine that the reload, addition, and immediate store affect onlye bits of the final store, and you can remove parts of it.
Pure speculation on my part that this is what happened, but all of these intermediate steps are performed all the time by modern optimizing compilers.
Maybe it chose this optimization due to register pressure, and not wanting to spill values that were still going to be used?
An I wouldn't be surprised if there's some weird x86 encoding that in some cases helps code size similarly, with different encoded instruction lengths allowing for different immediate sizes.
I always get a little anxious when looking at such code. But it seems to work well in practice?
"...each read of a single-word-sized or sub-word-sized memory location must observe a value actually written to that location"
It works in practice because it's a requirement of the implementation.
Right?
And of course only if you use safe Rust.
It is actually fairly tricky to get a value across threads like this. The simplest I could come up with is this:
struct UnsafeSync<T>(T);
unsafe impl<T> Sync for UnsafeSync<T> {}
fn main() {
let i = std::sync::Arc::new(UnsafeSync(std::cell::Cell::new(0)));
let thread_i = i.clone();
std::thread::spawn(move || {
thread_i.0.set(1);
});
eprintln!("i is {}", i.0.get());
}
https://play.rust-lang.org/?version=stable&mode=debug&editio...I could have used a Mutex then cast the `&Mutex` to `&mut Mutex` so that I could call `.get_mut()`. But this didn't seem simpler.
Because unsafe introduces conceptual complications and it seemed to be simpler to avoid that in an example.
I don't feel that the article take is useful or informative and amounts to gatekeeping ISA parallelism.