Code Optimization Techniques for Graphics Processing Units
hgpu.org
hgpu.org
You might also be interested in the work of a prof at the University of Alberta, Jose Nelson Amaral.
A Complete Descritpion of the UnPython and Jit4GPU Framework https://www.cs.ualberta.ca/system/files/tech_report/2011/Gar...
Jit4OpenCL: A Compiler from Python to OpenCL http://webdocs.cs.ualberta.ca/~amaral/thesis/XunhaoLiMSc.pdf
Cores are physically quite small; a lower clock rate reduces power issues) and I suspect that it would also increase yeild rates (perhaps by using a thicker structures with a finer process, e.g. 45nm on a 32nm process).
One barrier may be that practitioners have few techniques for such massive parallelism (catch 22). OTOH, it seems certain than manufacturers would have done their sums, and worked out they can deliver greater performance with their present number/clock tradeoff.
That NVidia monster has a die size of about 520 mm^2. At that size, there's already a lot of waste due to the fact that wafers are round and the chips are rectangular, and that can only be reduced by making physically smaller chips. (Rumor has it that by the time NVidia's 529mm^2 GF100 chip was originally supposed to launch, yields were bad enough that they were getting only about 2 usable chips per 300mm wafer. The cost of chip production has a pretty much linear relationship with the number of wafers processed, so that really hurt.)