http://www.nvidia.com/content/PDF/fermi_white_papers/D.Patte...
http://www.nvidia.com/content/PDF/fermi_white_papers/D.Patte...
I'm not sure that's impossible, it just seems very hard.
If nvidia manages to crack that nut then the only thing you'll still need to keep in mind is how big your cache footprint is (as on every other cpu with a cache) in order to maximize throughput.
That would definitely be a good thing.
I've spent in total about 2 months now (spread out over the last year) understanding how this whole GPGPU thing fits in with the rest of computing, it is much like a specialty tool. It is harder to master, more work to get it right once you have mastered it, subject to change on shorter notice than most other solutions (because of the close tie to the hardware) but if you need it, you need it bad and the pay-off is tremendous.