Moe style router?
My idea is more about creating recurrent paths in an otherwise all-forward network. The same number of weights would be loaded, we'd just be routing residuals differently.
[0] Usually multi-layer perceptron / linear portion weights - although maybe someone's tried attention head MOE?