One key idea here is to use a very large number of parameters (model weights), but only use some subset of the parameters on each example. The parameters are divided up into blocks called "experts", and then some subset of experts are used on any given input. Which subset is used is chosen by the model itself in a data-dependent manner. This can be thought of as letting the model specialize different experts to handle different situations.
The advantage, as they show, is that the model can train to a given level of performance much faster with a fixed amount of computing power compared to an architecture that uses all parameters on every step. This might be because it allows you to have a very large number of parameters that can store a lot more specialized information without incurring as much of a computational cost. Of course the downside is that you end up with a very large model that literally won't fit in a lot of environments.