I am pretty sure this is not correct. Every reordering effect I'm aware of is a core-local effect. That is, it happens in the core before (or at the moment) the data hits the L1D. It does not occur due to a weaker cache coherence system.
Having a cache coherence system which was itself weaker, allowing reorderings consistent with the memory model, makes barriers and "implicit barriers" like address dependencies very expensive, and there is little evidence this is the case.
Even in a hypothetical core which had a cache coherency system coupled to the memeory model, you aren't really avoiding any coherence traffic, just allowing certain reorderings such as satisfying requests out of order.