How much do you think it would really cost to develop an OpenCL equivalent of CuDNN (even a stripped down version, just fast)? I know AMD are struggling but we are talking about allocating a handful of talented engineers
Having C only wasn't a good idea. NVidia was quite clever in giving first class treatment to C++, Fortran and any compiler vendor that wished to target PTX.
Also the visual debugging tools are quite good.
Khronos apparently needed to be hit hard to realise that not everyone wants to be stuck with C for HPC in the 21st century.
Also although Apple is the creator of OpenCL, they don't seem to give much love to it.
Then you have Google caring about it's Renderscript dialect, which doesn't help to the overall uptake in OpenCL.
There isn't a monopoly, rather vendors that lacked the perception to appeal to the developers wanted to have as tooling and performance.
Anyone is free to go use OpenCL, use C or a language with a compiler with a C target, do printf debugging and feel free.
Are any vendors already doing SPIR support?
It's like saying MATLAB has a monopoly in academic research because so much of the code is written in it. That is slowly changing and moving over to Python now, which is great. Maybe OpenCL will get there someday, but I don't see it happening any time soon.
I would love it if AMD would care more about GPGPU, but they don't, and NVIDIA has little incentive to make their OpenCl drivers equal to their CUDA ones.
There's also Intel's MIC to consider now to, although that has a vastly different architecture to GPU. Again performance was similar between MIC and GPU in 2013[3], each performing better where their architecture was more suited, GPUs were capable of providing double the bandwidth for random access data.
In terms of AMD vs NVIDIA, I've not looked into it, I doubt AMD has anything to really compete with NVIDIAs current GPU accelerated compute lines. However again there was always that distinction (re bitcoin?) that AMD cards have better integer arithmetic and NVIDIA better float arithmetic.
Disclaimer: I use CUDA in my research, never tried OpenCL.
[1] http://arxiv.org/abs/1005.2581
[2] http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=604719...
Also, much of the speed gains for ML on NVidia hardware come from CuDNN - there is no equivalent for OpenCL or AMD hardware
AMD's Boltzmann initiative won't solve the lack of libraries.
CUDA had Fortran and C++ since day one and thanks to PTX was quite easy to add support for other languages.
Whereas OpenCL was stuck on "C only" model from Khronos, which forced everyone to use C or generate C code and be constrained to the device drivers.
This has been seen as such a big issue that SPIR and C++ SPIR got introduced with OpenCL 2.0.
Another very important one is debugging support. Last time I checked no one had visual tooling at the same level as NVidia's one.