Interesting, though the overhead of nanosleep vs. futex wait with non-NULL timeout should be dwarfed by the context switch. I can see where if you have extreme latency and/or throughput constraints on the writer side of the ring buffer, you couldn't afford the extra atomic read and occasional futex wake on the writer side to wake the consumer.
> but being written by only one process, ownership of cache lines never changes, so hardware locks are not engaged.
Which cache coherency protocol are you referring to where cache lines are "locked"? I'm aware that early multiprocessor systems locked the whole memory bus for atomic operations, but my understanding is that most modern processors use advanced variants on the MESI protocol[0]. In the MOESI variant, an atomic write (or any write, for that matter) would put the cache line in the O state in the writing core's cache. If we had to label cache line states as either "locked" or "unlocked", the M, O, and E states would be "locked", and S and I would be "unlocked".
Your knowledge of cache coherency protocols is probably much better than mine, so I'm probably just misinterpreting your shorthand jargon for "locked" cache lines. By "hardware locks are not engaged" do you just mean that the O state doesn't ping-pong between cores?
[0] http://www.rdrop.com/users/paulmck/scalability/paper/whymb.2...