For me, the place we'll eventually end up is obviously custom deep learning / evaluation chips that perform analogue operations using transistors in their linear regime (like how op-amps work). These chips would be programmed merely to express the tensor operation graph, essentially analogue tensor FPGAs.
This should bring multiple order-of-magnitude reductions in power consumption and increases in evaluation speed. And when you don't have a clock there might also be interesting ways of dealing with time in which you don't discretize and unroll, like one currently does with GRUs or LSTMs.