For me, AOT compilation is not about shipping binaries vs. source. The software I write is mostly used for internal research, so it's not going anywhere. It's difference between getting compilation errors from 'make' or having to pull them out of the OpenCL C API at runtime. Setting a breakpoint in an OpenCL kernel... I remember Intel's stuff basically working fine, but you had to pass a vendor-specific option to hint at the file path. In all these little ways, the workflow is just sucky. Much of it can be worked around, but I'm too lazy to write more application code to do all the chores that the CUDA toolchain takes care of already. I'm glad to hear that 2.0 is fixing this.
What's your alternative to templates for generic code, exactly? C macros? Scripted pre-processing that further screws up the already marginal tool support for debugging and profiling? Copy+paste? Templates are completely orthogonal to "complex computation" -- I just want to use a device function on different data types without run-time overhead.
On the topic of complex computation, I'm constantly surprised at the kind of features NV adds to CUDA and how well they actually work. I'm also surprised at the kinds of things people do on the hardware. If someone implements a high performance lock-free data structure on the GPU, you can't look at that and say oh, that's too complex, you shouldn't do that.
Also, until we're all working on computers that look like the PS4 with a unified global memory, there's a huge incentive to cram the awkward bits of your program onto the GPU any way you can, even if it drags a bit, because that's where the data is.
NVIDIA has really gone to some extremes. malloc in device functions. vtable support. Dynamic parallelism. Metal has none of this stuff.
<EDIT: Just now saw your other replies downthread. I'm leaving this comment because it reflects my personal experience and opinions, but don't feel like you need to repeat yourself to clarify your position re: templates, etc>