I think it is an interesting question. During the Alpha's lifetime as a non-hobbyist architecture, this decision was pretty much universally derided, but this was in the prehistoric eras of concurrent memory models, where people were just trying their best with a mix of C code, intrinsics, uses of `volatile` sprinkled around to hopefully disable optimizations, and inline assembly. When the C++11 memory model came around, they tried to integrate dependency ordering with `memory_order_consume`, and this famously failed, along with every attempt to fix it. I believe the plan is now for C++ (and later C?) to add special-cased RCU primitives.
The relaxation makes sense in the abstract. As evidenced from the `memory_order_consume` saga, compiler optimizations regularly violate dependency ordering anyways, so in the strictest sense you can't really rely on it. Of course, that doesn't stop people from "knowing" what their compilers will do in such a situation, but that strategy has become a worse one as the years have passed. I would feel better about the whole situation if there was a good greenfield design for a low-level PL that incorporates explicit dependency ordering.
I think it's less clear where the potential HW benefit is in a contemporary CPU. The obvious answer is value prediction, but any CPU performing value prediction has to deal with so many other microarchitectural conditions that can invalidate its speculation decisions that it's not clear this minor one is a huge burden. Of course, people who love TSO might make the same argument, i.e. that it's not a huge burden to snoop cache traffic and invalidate loads (although this is only the load half of TSO, not the store half). I know an Alpha architect who argued that Alpha was right for this decision. I even knew a Transmeta architect who argued for implementing sequential consistency in HW (as Transmeta and derived CPUs actually did).
In practice, microarchitectural structures have capacity/throughput limitations, and there are implementations and complexities that only come up in a real design, so everything needs to be evaluated in the context of a real project. I personally think the sweet spot falls to the weaker side of TSO, which also happens to be near the memory consistency model of the low-level languages we're using anyways.