The bottleneck actually is arithmetic. "GPUs have much higher ALU throughput since the GPU chip area is almost entirely ALU"
http://devblogs.nvidia.com/parallelforall/bidmach-machine-le...
Also on the horizon there is 3d chip manufacturing technology(3d-monolithic) ,with extremely large bandwidth between the two different layers of the chip,possibly being gpu + dram.