There are a couple potential reasons. Powerful GPUs accelerate research and iteration. Some state of the art problems have hit the limits of current theory and make up the deficit by building massive nets - but even there we already have multiple automatic pruning/optimization algorithms to shrink those nets so that they work with smaller resources.
Make no mistake, the field is advancing exponentially. The state of the art googlenet/inception that arguably kicked off the whole craze with image recognition are laughably obsolete now and easily outperformed by simpler nets.
MNIST was the gold standard for recognition problems just a couple years ago, and now it's considered a solved toy problem.