Since I've already provided a concrete example of what you asked for, I'm instead going to try to tackle head-on what I think is preventing you from understanding why the answers we are giving you are the answers to the questions you are asking.
C is not some thin translation layer for assembly, nor has it been for my entire lifetime (and I really wish people stopped teaching C as if it were merely "portable assembly"). At best, it is a description of a machine that does not, and will never, exist in practice, and the job of the compiler can be viewed as trying to emulate that machine using existing hardware. (This is still a somewhat poor description, but it does the job a lot better.) Trying to use hardware to understand how C works can be a fools' errand, because it's approaching the problem backwards. We don't try to describe hardware in C, rather, we try to think how C can be efficiently implemented--emulated--in hardware.
This comes to a very clear head when it comes to memory models. The abstract machine envisioned by C and C++ is, fundamentally, a sequentially consistent memory model. Now, no multiprocessor implements a sequentially consistent memory model, it's just not performant. The reason we can get away with such an unimplementable memory model detail is because of the data-race-free property: if you have release and acquire operations that obey certain properties, and your program is free of inter-thread dependencies that don't go through these operations (a data race, by definition) [1], then even weak memory ordering implementations are observationally indistinguishable from sequentially consistent implementations.
Because of the data-race-free model, the release and acquire barriers are fundamental to understanding the correctness of code. If they don't exist, then the code is definitionally wrong. And since the data-race-free model is nearly as old as I am, it's very well understood what hardware instructions need to be added to implement these barriers on every major platform. On some architectures, notably x86, the realization of these barriers is to do absolutely nothing; the hardware of loads and stores is sufficient to provide the guarantees these barriers require to preserve the illusion of sequential consistency.
Of course, barriers are still expensive on weaker processors, and people want to avoid them if possible. And some people noted that certain architectures--e.g., ARM and PPC--the hardware will guarantee the correct ordering of dependent loads and stores without a barrier. So the committee looked for a barrier whose realization on those architectures would be, like acquire is on x86, absolutely nothing. Thus we get https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2008/n26..., the first proposal for memory_order_consume. The first architecture mentioned is actually ARM, to motivate why control dependencies aren't included in the definition of memory_order_consume. Alpha is only mentioned, in the same breath as x86, in a note pointing out that it doesn't benefit from the proposed memory ordering. And if you read all of the subsequent papers on fixing memory_order_consume to make it usable [2], the entire motivation is not Alpha--where there is no meaningful difference from memory_order_acquire--but PPC, ARM, even Itanium.
Let me reiterate again because it's important. The entire raison d'être of memory_order_consume is to support hardware where its realization is to not insert a hardware barrier. The people who keep trying to fix it, even now, are--by their own words--motivated by a desire to make it implemented as such on hardware like ARM. That a barrier is needed for Alpha is incidental to the proposal; if the Alpha never existed, the same people would still introduce memory_order_consume for the purpose of introducing the requisite compiler-but-not-hardware-barrier.
> So we're back to "consume exists so you can write C++ code in 2023 to run on an architecture Compaq cancelled in 2003". Right?
Not at all. And if the words of the people who proposed this feature, talking about its applicability to most hardware other than Alpha, aren't enough to convince you otherwise, then I am truly at a loss.
[1] This, incidentally, is why relaxed atomics is such a specification mess: it's trying to specify the semantics of an intentional data race in a model that's fundamentally incompatible with data races.
[2] e.g., https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2014/n43...