Using atomic instructions is still generally considered lockless with the idea that since the lock is at the instruction level and thus progress can still be guaranteed it is "lockless". If this is true in practice depends on a whole host of factors.
I do not doubt majority of cases a mutex is faster/more appropriate. However the key there is looking at what m.lockslow() or unlockSlow() does which is called on lock contention. They become spinlocks that are not entirely dissimilar to my implementation.
The idea is that a lockless algorithm is trading a worse best case i.e m.lockSlow is not called for a better worse case. Again if this is true in this exact case, I am not sure as I have not benchmarked it against a naive mutex based ringbuffer.