My understanding and experience is that it is not always trivial to get linear training speedups with additional machines.
It can be sometimes hard to get any speedup at all.
The difference with gradient descent can be described simply as:
* Single-machine training: take more steps
* Distributed training: take fewer but more confident and accurate steps. This could allow you to take bigger steps (learning rate) as well, but there is a limit to this as well.
They are not equivalent processes and it is an area of active research how to get an equivalent result with distributed training.