A friend of mine that worked with supercomputers back in the 90s used to say that a supercomputer was a device that converted a computational problem to a communication problem. Today this is true for mainstream computing as well. There are still problems that can take advantage of the massive amount of computing available (Some mentioned Convolutional Neural Networks), but many tasks that used to be CPU bound is now effectively IO bound.
But don't believe me, don't believe a word I type. Go prove the above wrong by building a killer AlexNet implementation out of this chip. See slide 9 of this presentation for why I don't think you can do this (too much data transfer):
http://www.slideshare.net/embeddedvision/tradeoffs-in-implem...
If you can reduce the size of the computation units you can have one (or several) for each layer, making it possible to hardwire the transfer between layers
Yes, I'm still waiting for memristors.
http://devblogs.nvidia.com/parallelforall/bidmach-machine-le...
Also on the horizon there is 3d chip manufacturing technology(3d-monolithic) ,with extremely large bandwidth between the two different layers of the chip,possibly being gpu + dram.
The energy cost of transferring a single data word to a distance of 5mm on-chip is higher than the cost of a single FLOP (20 pico-Joules/bit). 5mm =~ the distance to L2 cache or another CPU core. The cost of transferring data off-chip (3D chip and/or RAM) is orders-of-magnitude higher, see graph.
[0] http://iwcse.phys.ntu.edu.tw/plenary/HorstSimon_IWCSE2013.pd...
What you run in to are problems feeding the beast data fast enough.