Yes, that’s how it works (pipeline parallelism)
Let's say the model has 50B parameters and 50 layers. That would mean about one billion values have to travel through the wifi for every generated token?
I wonder how much data that is in bytes and how long it takes to transfer them.