A few things:
- Details about GPT4 aren't very public. That's likely just a guess as to what's happening.
- 100T parameters isn't actually that hard to use on modern hardware. Getting it right with accelerators in the mix is a bit more challenging, but the assumption is exaggerated.
- An analogous idea in the literature is "dropout". So long as the network is trained with that same behavior it's usually beneficial. Dropping weights like the article describes is akin to severing connections in the graph of neuron connections. Traditional dropout ignores outputs at a given layer, which is akin to severing a whole neuron's worth of connections.
- Some intuition about the lack of warping is that dropout is equivalent to applying the dense transformation, dividing by that probability (to undo the skewed mean), and adding some noise. Training with dropout produces a network robust to noise (sort of like how tanh(100x) strongly reduces noise to push outputs toward a particular behavior -- compressing large domains into small ranges). Applying noise at inference time is atypical (usually the network is fully applied and then rescaled by that probability to have deterministic outputs fully using all weights), but it shouldn't yield drastically worse results.
- They'll likely get better results at inference time if, if they're averaging (might just do one sparse pass and call it a day) they do so at each layer. It helps reduce the chance of noise amplifying through the network and isn't any more expensive to compute (in any reasonable implementation it'd have lower communication costs, so should be a bit faster).