Yes, but the bulk of bandwidth should be eaten up by intermediate results rather than the input data e.g., for imagenet, 224x224x3 tensors are tiny compared to the intermediate activations that have to be moved around (and saved during training).
Hopefully you would also only copy your model over once.