For production loads? No one will run DS without expert parallelism so it’s not important for model to fit on one gpu.
Same for gpudirect (more useful for scale-out or training).
Everyone is focused on AI but these are two interesting techs trickling down from the NVIDIA tree, relatively "easy" to use there, which maybe exist in AMD world but I somehow missed the docs and APIs and demos on how to use them...