Regarding the performance, I'd be interested to hear more about your experiences. Because we compile to CUDA or OpenCL in the end, we cannot claim to be faster than what you could (in principle) write in hand. However, most of our benchmarks compare favorably with handwritten reference implementations, and the heavily optimizing compiler is able to write code that is tough to write by hand.
However, we're always looking for instances where we can do better. The goal is to be comparable to CUDA/OpenCL in as many cases as possible.