Need OpenCL support. Please don't contribute to maintaining CUDA monopoly.
I'm not a big fan of vendor "standards", but I have very limited sympathy for OpenCL here.
I think the best hope for portability is at the higher level programming API layer. For example TensorFlow is careful to make switching between CPU and GPU painless.
However the benchmarks for OpenCL look about 5x slower than CUDA