CPU inference is only a little slower. GPU's aren't good for a batch size of 1 and everything quantised.
The reason why GPUs seem to be the standard de facto is that they scale better, are more power efficient and are better supported by pytorch & co. Also, academia cares more about getting the best quality for their benchmarks, than about the performance and accessibility.