The holy grail would be a direct replacement backend that could be fed into TF, like CUDA.
ROCm has a direct replacement backend that can even take CUDA code (it's designed to be incredibly similar). It's called HIP. It's just that no one wants to support it. That is actually how TensorFlow on AMD works (mostly), and you can compile the latest stable release that way.
I can boot into Linux, and swap into Windows in 2 seconds with this setup. I have a dirty 20 line Bash script that deals with detaching the console, and passing the right things to the right place, but it all works.
ROCm on consumer cards does not work well. The tooling sucks. Massively. I don't understand why AMD doesn't have an extra team of 20 devs working just on the tooling.
Using DirectML with Windows Subsystem for Linux gives you better ML GPGPU support then AMDs native tooling.
That's a matter of opinion I suppose, but I don't personally find that passable.
>Using DirectML with Windows Subsystem for Linux gives you better ML GPGPU support then AMDs native tooling.
DirectML sucks even more than ROCm, IMO. Also, WSL sucks more than a normal VM.
What is a setup that would be passable then? I think a setup like the one that I have described [I believe] would be impossible with Hyper-V or ESXi (though, not that I have even attempted it with either).
Compare this to Nvidia or even Intel where I can run GPGPU API's (CUDA, oneAPI) on low-end, consumer-grade hardware.
The problem, ultimately, is that AMD does not even care.
I agree that AMD is dropping the ball.
CUDA support (or rather lack thereof, which isn't AMD's fault) plays a major role w.r.t. software support, but AMD's compute architecture isn't DL-focused either.
The MI250X compute part for example has a FP16 to FP32 ratio of 8:1 and a FP32 to FP64 ratio of 1:1; in other words it's an absolute beast at GPGPU compute.
The 6900XT on the other hand has a FP16 to FP32 ratio of just 2:1 and a FP32 to FP64 ratio of 1:16 (i.e. it's severely restricted at high precision workloads and OK at half-precision).
Comparing this to the specs of NVIDIA cards, they're still vastly superior on paper. The consumer versions of Ampere only get 1:1 (FP16:FP32) and 1:64(!!! FP32:FP64) respectively. But then again, NVIDIA cards feature dedicated "tensor cores", which have no equivalent on AMD consumer grade hardware.
The main selling point for NVIDIA, however, is software support and mindshare. They started to buy themselves into academia in the late 2000s by sponsoring labs and providing a vast ecosystem of software libraries for deep learning and GPGPU support in general. This not only helped kickstarting the deep learning revolution but also tied their hardware and brand name to GPGPU, which basically became synonymous with CUDA at that point.
Most other software focused exclusively on CUDA, so GPGPU on AMD cards is not well supported despite the potential.
Blender is a popular example for this. After discontinuing cross-vendor OpenCL support, which effectively limited GPU acceleration to NVIDIA cards, they added support for AMD cards in the latest version. Even then, only the latest generation of AMD cards is supported for some reason.
Other renderers like OctaneRender still only support CUDA. The situation is just as bleak in video editing software, were major companies like Adobe only support NVIDA and Intel (at least on Windows) or have poor OpenCL support (which users then blame on AMD of course).
There is some hope that software support for GPGPU without CUDA (maybe using Vulkan Compute?) will improve later this year with Intel (re-)entering the discrete desktop GPU market.