The accuracy would have to be significantly higher that any alternative to justify monopolizing that many hardware resources.
This kind of model could be used as the teacher in a distillation setup too though. Then faster training of the teacher is actually a huge benefit since it speeds up model development iteration cycles.
But even if it weren't practical to use in production in any sense, I'd argue there's value in doing the basic research of exploring design space of architectures in this way. This came out of a research team at Google. It may inspire and inform smaller, more practical architectures.
Part of the justification is the MoE sparsity by design means only a small part of the model will be activated by a given query. Don't think of it as a single giant 1t model, think of it as 50 small models which happen to share a glue layer at the input. So at deployment, you could, for example, keep the gating-layer in RAM and only pull the necessary sub-model off disk as necessary. Or you could shard the sub-models over 51 GPUs and feed the master gating-layer+GPU lots of queries, and each query will be dispatched to a different expert+GPU pair. This could easily be competitive with running a lot of dense models in parallel trying to keep up with the same load.