If you can just put the parameters into the multipliers in these layers, and leave them there while you cycle through the training data, you can use a lot less bandwidth to the compute hardware. Imagine a pipeline where you can feed through a billion tokens/second, each layer of the model persistently mapped to part of a grid of chips.
I suspect we're going to end up with a wildly different architecture before it's all over.