Can someone who understands this tell us what is the catch here?
Additionally, computation times grow quickly with higher resolution and you already need a high end GPU for this resolution to get a reasonably interactive response time.
You'll also need a favorably licensed pretrained model or a few 10000 training images and masks.
So all in all, I can't see any deal breakers, but I'd probably still use PatchMatch instead.
They say they used V100 but not how many, if they needed a large number then nevermind.
To that end, "reconstruct" is not so much reconstructing what was lost, but more "fills in holes" of photos. If you read the title expecting something to literally be reconstructed, you are likely to feel lied to.