The example isn't explained well in general, IMHO.
Suppose thread 1 executes the following code:
x = 1;
store_fence();
p = &x;
At a hardware level, it's guaranteed that the cache coherency traffic to update the value of p is going to happen after the traffic to update the value of x. So it's natural to assume that means that anyone who sees that p == &x must have to see that x == 1 as a result of this traffic. And for most architectures you'd be correct.
But you would be wrong on Alpha. Alpha has two cache banks, and there is no coordination between them (in the absence of memory barriers). So if p and x reside in different cache banks, it's possible for a thread to load the value of p (observing that it is &x), and then load the value of *p and fail to see the assignment of x = 1--if the cache bank that contains x is somewhat overloaded on processing the bus traffic, for example.
Incidentally, you can't actually take advantage of the opportunity to avoid the memory barrier in this scenario on everyone-but-Alpha in the C++11 memory model, because it turns out that figuring out how to specify the (hardware) data dependency matters for ordering purposes in a source language is a lot trickier than it might appear. memory_order_consume (added for this purpose) is lowered to memory_order_acquire in all compilers I'm aware of.