Safe Lock-free Primitives with iceoryx2's ByteAtomic
ekxide.io
ekxide.io
A sequence lock is a blocking lock. If the writer dies between the two increment operations, then the readers will spin forever waiting for the counter to become even again.
All the claims about something being lock-free and/or wait-free depend on a rational behavior of the writers.
If any writer acts crazy, progress becomes impossible regardless of what all others do, unless someone kills the rogue thread or process.
In practice, the algorithm from TFA is much more likely to guarantee progress than any of the algorithms that are theoretically proven to guarantee progress, because it has an extremely small overhead, while the alternatives are much more complex and they waste a lot of time.
Moreover, most wait-free algorithms guarantee progress only for the whole system, in the sense that one random thread will progress, but they do not guarantee anything for a given thread, which may be blocked forever or stuck in an infinite loop, if unlucky.
If the value changes frequently, it will get outdated quickly, but that has nothing to do with the synchronization mechanism used. And even if writes happen rarely, there is always a chance that the value you read will be outdated a nanosecond later.
If you have a bigger shared data structure, in which some other thread writes continuously, there exists absolutely no way to stop it and no way for any other thread to progress.
If a writer writes continuously the shared data, it is impossible for the other thread to make the copy that must be edited.
If the copy succeeds, then you are right that an updated version could be substituted to the original using an atomic operation on pointers.
But there is no way to guarantee that the first copy succeeds.
Of course, in practice RCU is used very frequently, because all the other threads are well behaved and access the shared data for a minimum time, so the copy will succeed in most cases.
But absolute guarantees are impossible inside an algorithm expressible in an abstract programming language. Only using functions of the operating system to detect and stop a misbehaving thread can solve all cases.
Or are you talking about a scenario where a rogue writer essentially randomly modifies the shared data structure instead of using the designated write() function? Well, in that case all bets are obviously off.
I agree that there is the risk for a writer to be halted in the middle of its critical section, which would stop all the other writers and readers.
My point is that there exists no solution that is risk free, because if a writer enters an infinite loop while writing the shared data, that will stop progress in any other algorithm, regardless if it is claimed to be wait-free.
There exists no method to stop such a writer, except an external intervention from the operating system, which would have to use an IPI (inter-processor interrupt) to halt that CPU core and then kill the offending thread.
In my opinion a great number of lock-free or wait-free algorithms, all of which are proposed based on the fear of what happens if a writer is halted in a critical section, are completely impractical, because their overhead is many times higher in comparison with using a lock for writers and using the method from TFA for readers.
With those algorithms, a lot of CPU time is wasted continuously to guard against an event that should never happen in bug-free operating systems and applications.
It is much more efficient to try to detect the lack of progress and do something about that only in the unlikely case when this happens.
Even if you have
while (true) { sharedData.writeWaitFree(randomData) }
all other threads will be able to continue. Whether the result will be of any value will depend on the use case.If, on the other hand, you mean that some threads will enter an infinite loop inside of a read or write operation, then you have a bug in your wait-free algorithm and all bets are off. But we would generally assume that the implementation is good and the erroneous behavior is external.
In that case no other thread can make progress.
I agree that this is a very unlikely case, but the case when a thread is halted inside the critical region can also appear only as a consequence of some bug, and such unlikely occurrences cannot justify wasting time at every access of shared data by using a too complicated wait-free algorithm.
No, it won't, not in a wait-free algorithm. For lock-free algorithms, yes, it's a matter of scheduling and stochastics.
But this is not the case for seqlocks. You don't need an infinite loop, you don't need to keep writing. Just stop the thread after it set the seqlock value to odd ("being modified"). Because it's not lock-free.
(I do agree that a lot of this is overblown and ill-applied; "lock-free" just sounds good and it's sufficiently available that people reach for it and end up overusing it. However, there are cases where it matters and is absolutely appropriate, and it also matters that we are able to have a conversation about these situations and conditions and use the terminology in a consistent manner.)
> using a lock for writers and using the method from TFA for readers
Case in point, I'm confused what you mean there, what do you mean with "method from TFA"? I don't see how anything in the article combines with a lock for writers.
Why is this a problem? Isn't the correct way to deal with a sequence lock failure to just retry? A torn read yes means you get undefined behavior as far as the result of your read, but you throw it all away and start again anyway so what is this solving?
As far as I know there is no reasonable compiler that wasn't specifically trying to add some kind of sanitizer that would emit anything other than plain memory reads, but it would be nice to not have to worry. But sanitizers are handy! I can imagine some sanitizer implementation forgoing extra internal synchronization that would only be needed in the case of program UB anyway.
As another poster has said, the high-level languages leave undefined what happens when you copy non-atomically data that is written concurrently, but in fact the computer cannot catch fire when you do that, and the only thing that can happen is that the data may have values that are invalid for its type, e.g. an integer that is defined to belong in a range may have a value outside that range.
A much more serious problem that is not mentioned in TFA is that on computers that do not use Intel/AMD CPUs, this algorithm needs write barriers and read barriers. The writer must use 2 write barriers, after incrementing the counter before accessing the shared data, and before incrementing the counter after finishing with the shared data. Similarly the reader needs read barriers after the first reading of the counter and before the final reading of the counter.
The description of the problem explicitly says that this is not happening. The data is being thrown away if the counter comparison fails. So the only undefined behavior that I can think they might be referring to is the act of loading the data itself, even though it is thrown away if it is wrong. Is that what they mean?
One of the key operations in these algorithms is a memory copy using core::ptr::copy. However, this results in undefined behavior if one process reads the data while another process writes to it concurrently. Even if our lock-free algorithm reliably detects such a race, iceoryx2 cannot depend on undefined behavior in a safety-critical system. This blog post introduces our solution: a byte-wise atomic wrapper that enables well-defined concurrent copy operations. It also shows how it can be used to implement a simple sequence lock.
I'm also missing any acquire/release barrier annotations in your code snippets. If you're using sequentially consistent accesses you might as well just single thread your code, performance wise.
Lastly, in almost all cases it's way more efficient and appropriate to shuffle things on the whole-object level, posting and retrieving pointers, and not poke around inside objects (especially on the byte level). Check how rare the use of seqlocks in the Linux kernel is, compared to other RCU primitives. (and regarding "appropriate", cf. top-level comment by danbruc https://news.ycombinator.com/item?id=49168283 )