Consume does not AFAICT act as a memory clobber or optimization barrier to the compiler.
Consume does not AFAICT act as a memory clobber or optimization barrier to the compiler.
This is your confusion. It is precisely an optimization barrier. Specifically, it's an acquire barrier, but only for certain data-dependent subsequent loads. Compilers universally treat it as a global acquire barrier because it turns out that it's too difficult to ensure that language-level data dependencies translate to hardware-level data dependencies.
On x86, the hardware never needs acquire barriers, so the relaxation to an acquire barrier has no effect. However, although the hardware doesn't need acquire barriers, the compiler still needs to know where those exist in the program so that it doesn't reorder the loads anyways. Meanwhile, on Alpha, the hardware still needs an acquire barrier in the case where there's a data dependency. So in both the ultra-weak and the rather strong memory models, translating memory_order_consume as memory_order_acquire has no actual difference to the hardware outcomes.
This is why memory_order_consume is really for ARM et al--the compiler still needs to know about the existence of the barrier, but the hardware doesn't need the barrier. And this is why there's still efforts going on to fix memory_order_consume; the people who work on ARM and the like want to get to the point where they can write these codes without having the compiler emit acquire barriers.
I don't understand what that means. Compiler barriers are about changing the ordering of instructions. "DEpendent" loads have defined ordering already, you can't issue a load before loading the pointer register you're dereferencing!
I have to ask again: can you give me an example from any toolchain and any architecture other than Alpha where a consume generates a detectable difference in the generated code?
if (i == 3) {
y = x[i];
}
The compiler will happily replace this code with: if (i == 3) {
y = x[3];
}
Now, the load of `x[i]` is no longer data-dependent on `i`. This is a real optimization that happens in compilers, and the Linux kernel (which relies on something akin to release/consume, although not directly dependent on the C11 memory model) has documentation specifically warning against using the results of the consume-like load in comparisons precisely because this optimization exists and has caused problems in the past.This is also why compilers have generally ignored memory_order_consume, by-the-by. Supporting consume 'properly' (i.e., other than strengthening it into a memory_order_acquire) requires suppressing optimizations that do not preserve data dependencies, but consume leaks out too much to the point that it's basically not possible to prove that the dependency isn't part of a consume chain, and the optimization is too beneficial to be worth turning off entirely.
(Edit: it's also true that this is a regime where we've seen security bugs due to optimization exactly like this. And that's true. But that's also not a memory ordering issue and it's, again, not resolvable by use of the C++ consume load feature).
So we're back to "consume exists so you can write C++ code in 2023 to run on an architecture Compaq cancelled in 2003". Right?
[1] Though it's true that my second phrasing was more ambiguous than my first.
C is not some thin translation layer for assembly, nor has it been for my entire lifetime (and I really wish people stopped teaching C as if it were merely "portable assembly"). At best, it is a description of a machine that does not, and will never, exist in practice, and the job of the compiler can be viewed as trying to emulate that machine using existing hardware. (This is still a somewhat poor description, but it does the job a lot better.) Trying to use hardware to understand how C works can be a fools' errand, because it's approaching the problem backwards. We don't try to describe hardware in C, rather, we try to think how C can be efficiently implemented--emulated--in hardware.
This comes to a very clear head when it comes to memory models. The abstract machine envisioned by C and C++ is, fundamentally, a sequentially consistent memory model. Now, no multiprocessor implements a sequentially consistent memory model, it's just not performant. The reason we can get away with such an unimplementable memory model detail is because of the data-race-free property: if you have release and acquire operations that obey certain properties, and your program is free of inter-thread dependencies that don't go through these operations (a data race, by definition) [1], then even weak memory ordering implementations are observationally indistinguishable from sequentially consistent implementations.
Because of the data-race-free model, the release and acquire barriers are fundamental to understanding the correctness of code. If they don't exist, then the code is definitionally wrong. And since the data-race-free model is nearly as old as I am, it's very well understood what hardware instructions need to be added to implement these barriers on every major platform. On some architectures, notably x86, the realization of these barriers is to do absolutely nothing; the hardware of loads and stores is sufficient to provide the guarantees these barriers require to preserve the illusion of sequential consistency.
Of course, barriers are still expensive on weaker processors, and people want to avoid them if possible. And some people noted that certain architectures--e.g., ARM and PPC--the hardware will guarantee the correct ordering of dependent loads and stores without a barrier. So the committee looked for a barrier whose realization on those architectures would be, like acquire is on x86, absolutely nothing. Thus we get https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2008/n26..., the first proposal for memory_order_consume. The first architecture mentioned is actually ARM, to motivate why control dependencies aren't included in the definition of memory_order_consume. Alpha is only mentioned, in the same breath as x86, in a note pointing out that it doesn't benefit from the proposed memory ordering. And if you read all of the subsequent papers on fixing memory_order_consume to make it usable [2], the entire motivation is not Alpha--where there is no meaningful difference from memory_order_acquire--but PPC, ARM, even Itanium.
Let me reiterate again because it's important. The entire raison d'être of memory_order_consume is to support hardware where its realization is to not insert a hardware barrier. The people who keep trying to fix it, even now, are--by their own words--motivated by a desire to make it implemented as such on hardware like ARM. That a barrier is needed for Alpha is incidental to the proposal; if the Alpha never existed, the same people would still introduce memory_order_consume for the purpose of introducing the requisite compiler-but-not-hardware-barrier.
> So we're back to "consume exists so you can write C++ code in 2023 to run on an architecture Compaq cancelled in 2003". Right?
Not at all. And if the words of the people who proposed this feature, talking about its applicability to most hardware other than Alpha, aren't enough to convince you otherwise, then I am truly at a loss.
[1] This, incidentally, is why relaxed atomics is such a specification mess: it's trying to specify the semantics of an intentional data race in a model that's fundamentally incompatible with data races.
[2] e.g., https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2014/n43...
You... really did not. I asked for code that is correct on x86 or ARM with a consume load and incorrect without.
> I'm instead going to try to tackle head-on what I think is preventing you from understanding
I think I'd understand it much better if you gave me the concrete example that I'm asking for and which you keep saying exists. Frankly at this point I think I'm done being patient and waiting for the explanation in the expectation I'm going to learn something, it's clear it's not going to arrive. Mostly: I don't believe you, I think it's you that are failing to understand the issue and not me, and I'll go on with my more comfortable hardware-level understanding of memory ordering and continue to ignore the pontifications of those in the language standard community. Sorry if that's frustrating.
int i = load(p, memory_order_consume);
if (i == 3) {
int y = x[i];
}
If you replace the memory_order_consume with a regular load, or a memory_order_relaxed, your code will be broken, even on x86 or ARM. The compiler can, and in this case likely will (yay constant propagation), rewrite the subsequent load in such a way that it is no longer data-dependent on the atomic load. It is not a compiler bug; if you try to report it as such, the compiler implementers will take one look at your problem report, and tell you that it is your code that is broken, use memory_order_consume. You will similarly get no sympathy from the C or C++ standards committees, for this is precisely why memory_order_consume was invented.> Mostly: I don't believe you, I think it's you that are failing to understand the issue and not me
As someone who works on a compiler, and who sits on language standards committees, when I am explaining to you how someone who works on a compiler and who sits on language standards committees will read the standard, it's not because I fail to understand how to do so. I do realize that the way the standards are read and interpreted does not come naturally to most, which is why I gave that long explanation. Perhaps you do not wish to engage with the specification on this level of understanding. That is fine--but a consequence is that you will not be able to properly understand the reason for certain features, like memory_order_consume.
At the very least, I hope you can agree to stop spreading misinformation like "the purpose of memory_order_consume is to support the Alpha processor."
int i = load(p, memory_order_consume);
if (i == 3) {
int y = load(x[i], memory_order_consume);
}
it seems like bad design to require ordering between plain loads and annotated loads. We want to identify the two locations that are important, so the compiler can reorder the others however it sees fit, without regard as to their placement w.r.t. those two.This is is catastrophic, if, instead of relying on an hardware barrier, we rely on a load-load dependency chain originating from i; rewriting x[i] into x[3] break the chain and now, while the compiler might still not reorder the access, the hardware very much can.
If its name is DEC Alpha? ;)
Of course "do what gcc x.y does" is not a good spec so the standard tried to specify exactly the behavior of consume is, but it failed to be implementable.
Right. Solely so it can support DEC Alpha! I'm not saying data dependence has never been an ordering constraint, I'm saying it was only an ordering constraint on an ancient platform that thus doesn't belong in the standard.
And by extension I'm saying that you people talking about standards instead of hardware are needlessly (frankly: deliberately and obfuscatorily) obscuring that clear truth.
You must use one of the rcu_dereference() family of primitives
to load an RCU-protected pointer, otherwise CONFIG_PROVE_RCU
will complain. Worse yet, your code can see random memory-corruption
bugs due to games that compilers and DEC Alpha can play.
Without one of the rcu_dereference() primitives, compilers
can reload the value, and won't your code have fun with two
different values for a single pointer! Without rcu_dereference(),
DEC Alpha can load a pointer, dereference that pointer, and
return data preceding initialization that preceded the store of
the pointer.
The whole document then goes in great details on the all the ways the compiler can screw you over.If the kernel were to drop Alpha support tomorrow, rcu_dereference and all the warnings would still be relevant.
Similarly, if only x86, SPARC and Alpha existed, memory order relaxed, acquire/release and seq-cst would be sufficient. consume only exists to cater to architectures with memory models stronger than Alpha but weaker than TSO.
[1]https://www.kernel.org/doc/Documentation/RCU/rcu_dereference...
Just because the simplest translation to assembly yields a data dependency on ARM et al doesn't mean the compiler is obligated to honor that.
The focus on standards is precisely because that's the only contract the compiler gives you. If you write code expecting to also get your target architecture's guarantees for free, you're writing software that is at best only incidentally correct until the next compiler update.
I look around this thread and I don't see a group of people trolling you - I genuinely think you're missing the point.
You keep saying that somehow this isn't about alpha, but it's clear this is about alpha, because alpha needs a barrier there and volatile can't do barriers. You just... won't admit it. And per my point WAY upthread, I contend you won't admit it because you view the "standard" as the important thing and not the hardware.
Even if volatile was absolutely all you needed, why would the committee build a memory model with a nice API and then force you to sprinkle in volatile? It makes absolutely no sense.
You can build yourself an access API if you have volatile. E.g.
#define volatile_lvalue(TYPE, LOC) (*(volatile TYPE *) &(LOC))
With typeof or _Generic, we can eliminate the TYPE argument.Then you can do:
if (volatile_lvalue(int, i) == 3) {
y = x[i]; // don't care if i replaced with 3 here
}
The user of volatile_lvalue doesn't have to declare anything volatile.In practice the volatile cast might prevent current GCCs from performing constant propagation, so this exact code sequence might work, but it is extremely fragile and can break for some variations:
if (int i2 = volatile_lvalue(int, i); i2 == 3) {
y = x[i2]; // i2 is replaced, no data dependency
}
or if (volatile_lvalue(int, i)) {
y = x[i-i]; // data dependency becomes a control dependency
}
You also have to be careful to what you do with y and with any value derived form it.As currently there is no way to efficiently implement consume you have to do with what you have, so for example the Linux kernel does use volatile casts and compilation barriers to implement LOAD_ONCE, but then has a large manual that documents what operations are currently considered safe and what are known to break.
That's hardly a specification that can be put in a standard and you can't really can reason about the semantics of it. Linus himself has considered multiple time if it is worth the hassle and whether the kernel should just switch to acquire everywhere now that arm and power have relevant cheap barriers.
volatile int i;
volatile int *x;
if (i == 3) {
y = x[i];
}
Even if this becomes x[3], volatile should work well enough to prevent the reordering: the access to x[] shouldn't precede the access to i. Now volatile might cause two accesses to i. That's my problem, which I could try to defeat with caching: int local_i = i;
if (local_i == 3)
y = x[3]; // Let's do this ourselves myself
Still, everything should be cool; the access to i and to x[3] should be in that order thanks to volatile. We have one access to i.If, in spite of the compiler obeying our volatile, the processor and memory hardware is rordering our accesses, then we need some fancy primitive, because special instructions have to be used.
The fancy primitive should exist only for the hardware problem, not for the compiler's ordering of instructions.
[1] I'm also ignoring the fact that volatile really doesn't have the right semantics in C++11/C11.
So let's not fix that, but invent new cruft.
C defines volatile for several purposes of its own, and that's it.
- In a function that uses setjmp, volatile must be used to mark variables that are modified after the context is saved with setjmp but before control returns with longjmp. Otherwise there is a problem, like the longjmp spuriously restoring the original values of those variables (due to those variables having been pulled into the machine state being saved and restored, like registers).
- When an asynchronous signal handler (e.g. for SIGINT) sets a global variable which is checked by the interrupted code, that variable has to be "volatile sig_atomic_t".
Any usefulness of volatile for actual concurrency is courtesy of the compiler, whether empirical or documented.
As an extension, MSVC did for a time add atomic and acquire/release semantics to volatile. This was deprecated because it broke too much stuff.
edit: after your edit, I'm not sure if you arguing for volatile-for-atomic or not.
Indeed I've worked with hardware where it was more obvious how to lower an atomic load or store than it was to lower a volatile load or store, because the technical definition of volatile really doesn't comport with the practical meaning of "don't optimize this" (crazy hardware does things like that).
I.e. the quirks of Alpha have no bearing on the existence of consume. A different annotation between relaxed and acquire would still needed to specifically capture the semantics of data dependencies (that's the theory at least, in practice even consume and the [[carries_dependency]] annotations are not enough).