Well the batch size also helps with making training more stable. So a larger batch size computes the gradient across more examples so it ends up being more stable than only using a batch size of day one. A batch size of one is only fixing the error on one example rather than many.
a larger Batch size can help keep the gpu fed but usually that’s not a problem with these larger models.