> buffer.write.store(buffer.write.load() + write_len)
This is... an atomic load, followed by an add, followed atomic store.
A TRUE lock-free queue would be buffer.write.AtomicAdd(write_len), (which is a singular, atomic add).
There are many other issues here, but this is the most egregious issue I was able to find. The whole thing doesn't work, its completely non-safe and incorrect from a concurrency point of view. Once this particular race condition is solved, there's at least 3 or 4 others that I was able to find that also need to be solved.
EDIT: Here's my counter-example
write = 100 (at the start)
write_len = 10 for both threads.
|-----------------------------------|
| Thread 1 | Thread 2 |
|-----------------------------------|
| write.load (100)| write.load (100)|
| 100+write_len | |
| write.store(110)| 100+write_len |
| | write.store(110)|
|-----------------------------------|
write = 110 after the two threads "added" 10 bytes each
Two items of size 10 were written to the queue, but only +10 bytes happened to the queue. The implementation is completely busted and broken. Just because you're using atomics doesn't mean that you've created an atomic transaction.It takes great effort and study to actually build atomic transactions out of atomic parts. I would argue that this thread should be a lesson in how easy it is to get multithreaded programming dead wrong.
---------
For a talk that actually gets these details right... I have made a recent submission: https://news.ycombinator.com/item?id=20096907 . Mr. Pikus breaks down how atomics and lock-free programming needs to be done. It takes him roughly 3.5 hours to describe a lock-free concurrent queue. He's not messing around either: its a dense talk on a difficult subject.
Yeah, its not easy. But this is NOT an easy subject by any stretch of the imagination.