To quote the paper:
> Our model is supervised with paired input and target views of a scene (along with their camera poses)... The model then needs just a single image at test time.
Correct me if I’m wrong, but: Given a novel scene, it seems the model must be retrained on multiple images of that scene? It seems disingenuous, then, to say it works from a single image. No doubt the interpolation is state of the art, but the title seems misleadingly magical.