Measuring the Algorithmic Efficiency of Neural Networks
arxiv.org
arxiv.org
Things like grouped convolutions were invented for AlexNet as a practical engineering step because of limited GPU memory, but ended up giving nice cost/accuracy trade-off choices.
Perhaps algorithms will move too fast for dedicated hardware to be worth it, but there will be primitives that should be relevant for a while that can be integrated into whatever hardware we use - see Nvidia Tensor Cores, which also include things like sparsity support.
The larger story today is Azalia Mirhoseini's paper in Nature: Deep RL for next gen chip floorplanning in AI accelerators. Future doubling rates of two hours rather than two years on the horizon