My view is CUDA has already won, and everyone else needs to get over it. Even clang supports PTX now, which is a reasonably device agnostic representation, albeit controlled mostly by Nvidia. Perhaps intel will introduce their own extensions to this ISA.
Even if my precompiled CUDA application could run on Intel GPUs at 50% of the throughput, I'd be happy if I could later tweak and recompile it to get the full benefits from their hardware.