The paper has a fundamental flaw, which can be seen if trying to reverse the reasoning to look at the past, rather than trying to predict the future.
If you used today's level of computation with 2012 era models, then you wouldn't have the error rates of today's models, you would have much higher, much worse errors.
The single biggest bottleneck in deep learning is not computation, it's the ingenuity used to devise new structures and new optimizations that allow for scaling.
For a given structure, with given techniques, you can throw more computation at a problem to decrease error rate, but the gains scale poorly with cost. Cost improvements only come with new techniques and structures.