This enables you to run larger models
than you would be able to on any single
device.
No further explanation on how this is supposed to work?If some layers of the neural network are on deviceA and some layers are on deviceB, wouldn't that mean that for every token generated, all output data from the last layer on deviceA have to be transferred to deviceB?