Traditionally, you would use a acquire, L-LS, barrier on entry to a critical section to order any memory accesses in the critical section after the load of the atomic operation completes. Without that, the critical section may speculatively execute on your core while somebody else holds the lock and those changes would be confirmed if the owner releases before your delayed acquiring load completes.
For this paper to be true, either their synchronization primitives do not have acquire barriers, which is wrong to start with, or the acquire barrier must not be stalling later memory accesses which is how I have seen it portrayed. How would that even work correctly on normal code?
To preserve the L-L ordering implicit in the L-LS ordering, you need to guarantee that any loads after the acquiring load will see all stores explicitly ordered before the releasing store. You would need to guarantee that any speculative load must at least see the value the memory would have immediately preceding the release barrier on the other core even if you happened to load that value arbitrarily long before the release barrier on the other core. I guess you could track every single speculative access after a L-* barrier and their dependency chain and then invalidate and re-compute as you see stores until you resolve all pertinent operations prior to your L-* barrier (such as the acquiring load in this case)? If the chip people are actually doing that, that seems... wild.