I updated the author's benchmark:
https://gcc.godbolt.org/z/bxEzerjc8While updating it, I noticed a couple of things:
1. the Makefile doesn't specify -O, which seems to imply -O0, unreasonably flattering the 1-memcpy case because there's only 1 call to a generic memcpy rather than 2. With -O3 the memcpys get inlined
2. make the message size more annoying. The queue size was a multiple of the message size, so the 2-memcpy case was never exercised - and -O3 could see this
(Change 2 actually slightly cancels out change 1, and now the 2-memcpy case is terrible again...)
Who even knows what sort of system the Compiler Explorer executor runs the programs on, but this is what I got from running it, and the relative speeds might be representative:
%: 0.033716 microseconds per write
&: 0.031772 microseconds per write
slow: 0.009715 microseconds per write
fast: 0.001035 microseconds per write