Last time I checked, Stable Diffusion was much slower on rocm than it was on an alternative stack (SHARK's IREE/MLIR stable diffusion demo), and it was much slower than comparable Nvidia GPUs. But that was pre torch 2.0.
And none of big SD repos are using torch.compile for inference yet... I have it hacked into Automatic and VoltaML, but its really finicky.