(and let's hope it goes that far down), that would make things a lot easier as well.
Unified address space I assume you mean across multiple GPUs ? Global memory cache is a double edged sword, that eats in to the transistor budget at a very rapid pace, effectively you already have a cache, you just have to fill it yourself.
GPU programming is definitely a step back in the ease with which you can write programs, but if your problem maps well on to a GPU the speed increases are simply astounding. What would have taken you a cluster with 100 boxes now sits under your desk and consumes 250 watt tops. That's really very impressive.
The way intel seems to edge in to gpu territory and nvidia into cpu territory will make for some interesting stuff happening in the next couple of years.