GhostRace: Exploiting and mitigating speculative race conditions
vusec.net
vusec.net
Traditionally, you would use a acquire, L-LS, barrier on entry to a critical section to order any memory accesses in the critical section after the load of the atomic operation completes. Without that, the critical section may speculatively execute on your core while somebody else holds the lock and those changes would be confirmed if the owner releases before your delayed acquiring load completes.
For this paper to be true, either their synchronization primitives do not have acquire barriers, which is wrong to start with, or the acquire barrier must not be stalling later memory accesses which is how I have seen it portrayed. How would that even work correctly on normal code?
To preserve the L-L ordering implicit in the L-LS ordering, you need to guarantee that any loads after the acquiring load will see all stores explicitly ordered before the releasing store. You would need to guarantee that any speculative load must at least see the value the memory would have immediately preceding the release barrier on the other core even if you happened to load that value arbitrarily long before the release barrier on the other core. I guess you could track every single speculative access after a L-* barrier and their dependency chain and then invalidate and re-compute as you see stores until you resolve all pertinent operations prior to your L-* barrier (such as the acquiring load in this case)? If the chip people are actually doing that, that seems... wild.
If I remember correctly, x86 basically gives you everything except S-L automatically. That means a traditional acquire, L-LS, and traditional release, LS-S, barrier should be "free" and require no explicit instructions. I forgot about that, so it makes sense why they would not mention the presence of explicit barriers. But, that does not change the fact that you need that ordering to guarantee correctness, it is just built-in and automatic.
The oddity here is that speculative loads across a L-LS ordering requires some pretty deep black magic to be going on at the hardware level. I have no idea how they could safely reorder stores across a L-S ordering and still maintain "as-if" the store happens after the load, and even safely reordering loads across a L-L ordering and maintaining the "as-if" seems pretty arcane. The fact that the paper is claiming their PoC works and thus the reordering must be happening is wild. Either the processor is not correctly implementing the declared memory ordering semantics, or they are doing some real reordering magic.
The basic way it works is that pretty much any reordered is allowed, speculatively, even if it violates the explicit or implicit barriers, but then the operations are tracked in a memory order buffer (MOB) until retirement which can detect whether the reordering was detectable and if so flush the pipeline (so-called "memory order nuke", visible in performance counters).
> or the acquire barrier must not be stalling later memory accesses
That is my impression. But I am really confused about this stuff in general.
The demonstrated exploit strategy is pretty cool.
5 percent is not nothing but seems like a worthwhile investment here. Really cool exploit strategy btw.
But isn’t this like the 4th or 5th mitigation by now?
Imagine paying extra for a 4ghz CPU only to find it operates no better than 3ghz on a good day…
I'll tell people it's a punny triple entendre, leading people to believe I am a wall street hacking guru hit-man.