Both mindshare and market share (outside of super computers) are just overwhelming at this point. Their market share in the consumer market is ~86% as of this quarter. The data centre market is quite fragmented when it comes to accelerators, but in AI training, NVIDIA is still the market leader.
> Game engines/engineers seem to be able to be productive on a wide variety of GPU hardware, why do so many ML libraries just not provide any support at all?
That's a different kettle of fish. Game engines rely on the graphics driver's implementations of low- and mid level APIs like Vulkan or Direct3D.
The brunt of the work is also often performed by the middle-ware (mostly Unreal Engine and Unity or in-house engines like Frostbite) that had been in development for decades; with most games focusing on high level optimisation wrt. the middle-ware used.
ML-frameworks on the other hand need to optimise compute kernels as well as data flow between host CPU and accelerator (e.g. GPU). This involves hand-tuning algorithms to best match specific GPU architectures, while most shaders used in games are basically the same across all GPU vendors and it's the vendors themselves who do the fine tuning and per-game optimisations in their graphics drivers (hence the obscene sizes of GPU drivers these days).
While that's a good enough approach for games, it's simply not possible to do the same for ML-models. There's just too much flexibility (no API that dictates which calls do what, when, and how) to make general ML-optimisations at the driver level.
> How come every ML library doesn't at least have an OpenCL fallback or the like?
OpenCL is horrible to work with and stopped being properly supported by vendors (e.g. newer versions are rarely being implemented and optimised). The difference between CUDA and OpenCL from an implementor's perspective is that CUDA works seamlessly with surrounding C++ code and compute kernels can be embedded in the host CPU code base. OpenCL on the other hand is modelled after the ancient OpenGL 2.x paradigm and requires tedious setup and careful integration (checking capabilities and all that jazz). OpenCL is basically dead at this point.
There are alternatives to CUDA, but most frameworks rely heavily on the highly optimised libraries NVIDIA ships (e.g. CuDNN) and don't have the resources to implement the functionality themselves. Some hardware vendors offer proprietary backends, like Apple or Intel and you just have to wait for them to catch up. AMD has ROCm, but that's more of a drop-in replacement that aims at running CUDA code on AMD cards.