A good spinlock, as Linux uses internally for example, is much more complicated. But, if you know there are only two relevant threads, you can get decent performance with a simple spinlock.
However, do not use AtomicExchange in a loop like this. When you are waiting, you always want to wait with the target cache line in the shared state, and AtomicExchange will force it to be exclusive. Performance will go down the drain. Instead, do something like:
if (value looks good)
AtomicExchange;
Loop again if you lost the race;
else
Loop again;