I agree with kbumsik here. AMD only has themselves to blame. They have great hardware and fantastic theoretical benchmarks. Heck, even their SGEMMs are really fast and in line with the 15TFlops of FP32 on the VEGA 64s that we've benchmarked. However, it comes down to software ecosystem and optimizations for common deep learning subroutines.
MIOpen[1] is a step in this direction but still causes the VEGA 64 + MIOpen to be 60% of the performance of a 1080 Ti + CuDNN based on benchmarks we've conducted internally at Lambda. Let that soak in for a second: the VEGA 64 (15TFLOPS theoretical peak) is 0.6x of a 1080 Ti (11.3TFLOPS theoretical peak). MIOpen is very far behind CuDNN.
Lisa Su, if you're reading this, please give the ROCm team more budget!