Can someone translate this post into english for those of us not imbued with AI knowledge?
The advantage, as they show, is that the model can train to a given level of performance much faster with a fixed amount of computing power compared to an architecture that uses all parameters on every step. This might be because it allows you to have a very large number of parameters that can store a lot more specialized information without incurring as much of a computational cost. Of course the downside is that you end up with a very large model that literally won't fit in a lot of environments.
Researchers at Google's scale prefer a single model where you throw all your data in a single bin and get perfect performance out, no tweaking and no pesky humans required.