I worked on a research project that set out to answer that question:
https://arxiv.org/abs/1811.03600
We looked at a bunch of different model architectures and datasets and found that you can get speedups from larger batch sizes up until a certain point, but that point is different for different datasets and architectures. The range of good hyperparameters is also narrower for larger batch sizes, which makes tuning harder.