atomic.AddInt32 feels a lot like locking to me. You end up with a write barrier and all the associated consequences, so you can kill throughput (in a way that is more difficult to notice in the profiler, though you'll probably notice it when it's a problem if you ask pprof for a line-by-line profile).
sync.Mutex.Lock is implemented like this:
func (m *Mutex) Lock() {
// Fast path: grab unlocked mutex.
if atomic.CompareAndSwapInt32(&m.state, 0, mutexLocked) {
if race.Enabled {
race.Acquire(unsafe.Pointer(m))
}
return
}
// Slow path (outlined so that the fast path can be inlined)
m.lockSlow()
}
lockSlow is, of course, a bit more complicated.