To observe the "epoch counters" between threads (which can run on different cores), don't you still need memory fences/barriers to make the epochs visible to other threads - which I found is a significant part of the cost of locking.
Perhaps that accounts for "only" 25m reads/s, although I'm surprised that mutexes are sooo much slower.