if your data exceeds memory then a GPU is worthless. distributed makes it scalable
Furthermore, almost any technique you use to distribute and scale training will work just as well regardless of whether the computations are happening on CPUs or GPUs.
http://research.google.com/archive/large_deep_networks_nips2...