This is a commonly known problem called false sharing. https://en.wikipedia.org/wiki/False_sharing
To be fair: I don't know how these particular atomics work, maybe they're all on their own cache lines, and 192 bytes are used up for those three atomics so that things are done properly (unlikely, but possible). But atomics in my experience usually assume the "external" programmer figures out the buffering issue to prevent false sharing.
2. These are atomics declared to be SeqCst memory ordering. That means the L1 (and probably L2) caches are flushed every time you read or write to them on ARM systems. (x86 has stricter guarantees and almost "natively" implements SeqCst, so you won't have as many flushes on x86). The cache must be flushed on EVERY operation declared SeqCst. Its absurdly slow, the absolute slowest memory model you can have.
3. Well-written spinlocks are absurdly fast. Benchmark them if you don't believe me. The assembly in x86 is
Lock:
while (AtomicExchange(spinlock, true) != false) hyperthread_yield(); // loop until you grab false.
acquire_barrier(); // Prevent reordering of memory
Unlock:
release_barrier(); // Prevent reordering of memory. NOP on x86 but important in ARM / POWER9
spinlock=false; // Don't even need an atomic here actually
Literally one assembly instruction to unlock, 3-assembly instructions to lock (the atomicexchange + loop) + hyperthread_yield if you wanna play nice with your siblings (reduces performance for you, but overall makes the whole system faster).From a cache-perspective, the spinlock will play nice with L1 cache (if a thread unlocks, then locks again before a new thread comes in, it all stays local to L1 and no cross-thread communication occurs).
So pretty much every aspect of the hypothetical spinlock implementation is faster in my mind. I don't really need to benchmark to see which one is faster.
EDIT: At bare minimum: because the Spinlock will be properly written using Acquire / Release half-barriers, they will be ABSOLUTELY faster than SeqCst Atomics. You'll need to use Acquire/Release Atomics if you want to "keep up" with the Acquire/Release based Spinlocks.