My gut says that the quality could be rapidly improved without changing the underlying design at all.
The real issue with this, I think, is that motion capture for humans is already widely available and provides much higher fidelity and control than text. Unless I'm misreading the paper badly, this model was trained on exactly such data. Blending between multiple animations through motion capture is also well-understood.
So while the results are impressive, the practical gains seem very marginal. I think perhaps that the equivalent to "inpainting" (as mentioned in the text) and "style transfer" would be the big gain here? If we could use this to retarget animations to different body plans (child, adult, space monster) quickly, or for smarter interpolation between human-authored keyframes, I could see that being a much-desired tool.