You bring up some good points, and have forced me to consult documentation to double-check a lot of intricate details.
I'm using the following as a reference: https://static.docs.arm.com/100941/0100/armv8_a_memory_syste...
> Fences for sequential consistency are certainly heavy on ARM and it will show up as poor performance, I'd be very surprised if any aarch64 processor flushed any cache on LDAR/STAR.
I agree. But LDAR / STLA are weaker than full sequential consistency. They're only acquire/release consistency.
So its the DMB ISH that I'm curious about. If the memory is never flushed out of L1 cache, how can you possibly establish a total sequential ordering with other threads?
EDIT: I had a bad example here originally. Lets use the SeqCst example from here instead: https://en.cppreference.com/w/cpp/atomic/memory_order#Sequen...
std::atomic<bool> x = {false};
std::atomic<bool> y = {false};
std::atomic<int> z = {0};
void write_x()
{
x.store(true, std::memory_order_seq_cst);
}
void write_y()
{
y.store(true, std::memory_order_seq_cst);
}
void read_x_then_y()
{
while (!x.load(std::memory_order_seq_cst))
;
if (y.load(std::memory_order_seq_cst)) {
++z;
}
}
void read_y_then_x()
{
while (!y.load(std::memory_order_seq_cst))
;
if (x.load(std::memory_order_seq_cst)) {
++z;
}
}
int main()
{
std::thread a(write_x);
std::thread b(write_y);
std::thread c(read_x_then_y);
std::thread d(read_y_then_x);
a.join(); b.join(); c.join(); d.join();
assert(z.load() != 0); // will never happen
}
Variable x is being written to by "write_x". While variable y is being written to by "write_y" thread. Assume they're on different cache-lines.
* Thread C may see "X=True THEN Thread C THEN Y=True" (Formally: Thread A happened before Thread C happened before Thread B). In this case, Z=0.
* Thread D may see instead "Y=True THEN Thread D THEN X=True" (Formally: Thread B happened before Thread D happened before Thread A). In this case, Z=0.
This apparent contradiction is possible in Acquire-Release consistency, but not possible in Sequential Consistency. LDRA and STLR provide only acquire-release consistency.
So now lets ask: how do you implement sequential consistency?
There's no way for Thread A or B to do things correctly from their side. Thread A independently writes "X" in its own leisure, not knowing that Thread "B" is writing a "related" variable somewhere else called Y. There is no opportunity for A or B to coordinate with each other: they have no idea the other thread exists.
---------
To solve this problem, I imagine that an MESI cache will simply invalidate the cache (which "forcibly ejects" the data out of the core running Thread A and Thread B). Whatever order Thread A and Thread B create (X then Y... or Y then X) will be resolved in L3 cache or DDR4 RAM.
Now... perhaps there's a modern optimization that doesn't require flushing the cache? But that's my mental model of sequential-consistency. Its slow, but that's the only way I'm aware of that establishes the strong SeqCst guarantee.
---------
However, you are right to challenge me, because I don't know. If it makes you feel better, I'll "retreat" to the more easily defended statement: "Load/Store Buffers of the CPU Core need to be flushed every time you SeqCst load/store on ARM systems" (which should effectively kill out-of-order execution in the core).
The question I have is how does L3 cache perceive a total-ordering without the L1 cache flushing? I guess you've given me the thought that its possible for it to happen, but I'm not seeing how it would be done.