However, the tooling and installation procedure is still rough around the edges and the documentation is lacking.
Last time I tried to upgrade the rocm things and the drivers, and bumped up against an issue that one of the packages tried to write to /data/jenkins_workspace during installation. We have /data as an NFS mount and is not writable even for the root user, so the installation broke down. Worked around it, but there are other issues as well. The order of packages you install can matter as well, I had to purge packages and reinstall them a few times. It worked with TF 1.14.1 but not with 1.14.0 if I recall correctly. The first few iterations are very slow, it seems it needs to find the optimal conv implementation for all kernel sizes separately.
Also, the card does not have anything like NVIDIA's tensor cores for half-precision (float16/FP16) computation. It can do FP16 but won't speed up and delivers worse results (in terms of prediction accuracy) than NVIDIA cards running FP16. There is no prediction accuracy difference in FP32.
And some other small issues here and there. So be warned and test it yourself on one card before investing in big multi-GPU machines.
But I do wish AMD entering the field will benefit us through increased competition in the long run. But at the moment ROCm seems like just a side project for a small team in AMD and they can't yet deliver the streamlined experience we're used to from CUDA.