My guess would be that the videos processed in their work are quite small to limit computation time
I imagine this would be much less convincing in higher resolution.
The Paper includes some higher resolution examples which are quite convincing
The paper is a PDF file, is there another video component I'm missing?
It's already not very convincing if you look closer. For instance, look at the way hair behaves (or rather, how it doesn't). Or notice the artifacts around shoes.
You cant, its a hard rule of image/video manipulation papers, results must be 64x64 pixel or smaller.