Worldsheet: Wrapping the World in a 3D Sheet, View Synthesis from a Single Image
worldsheet.github.io
worldsheet.github.io
> Our model is supervised with paired input and target views of a scene (along with their camera poses)... The model then needs just a single image at test time.
Correct me if I’m wrong, but: Given a novel scene, it seems the model must be retrained on multiple images of that scene? It seems disingenuous, then, to say it works from a single image. No doubt the interpolation is state of the art, but the title seems misleadingly magical.
Detecting Invisible People
Complex scenes obviously take longer, but results can be extremely convincing for a limited range of view angles.
I learned the technique from this video: https://youtu.be/RSC40B8kHT8
That's good enough for sideways views, but fails if you try to look behind objects, because in their approach, there's no "behind", just a big, stretchy sheet covering everything.
A model with more knowledge about the world would be able to predict that a tree trunk is roughly cylindrical and not connected to the background.
This paper does a good job of guessing what should go in a scene when given depths and shapes https://www.youtube.com/watch?v=u4HpryLU-VI&ab_channel=TwoMi...
A combination of the two would be wild
Generating new views from an rgb+depth image is relatively straightforward and I'd expect any reasonable implementation to work pretty well.
They just need to build it into their viewer, which I also wish they would do.
Imagine a situation where you're standing on a street corner with a mailbox or something standing on the corner. You move to the next photo position, which takes you past the mailbox. The smearing and warping you are seeing is from the continous mesh they've created from the Lidar data not matching the real world mesh in areas that were occluded from the original point of view. You get a moment of seeing "behind" the mailbox, but there is no data for what is behind the mailbox, so it all gets interpolated from the data that is known in the surrounding visual area. The mailbox becomes a rectangular prism that extends all the way from the mailbox's location through to the intersection with the ground that we can see behind the mailbox.
These are just the problems with single-image mesh recreation. You can't really get around them without some form of inference of the data that doesn't exist. You even see it in Facebook's images in the linked article, if you look closely. They try to cut the videos early so you don't see it, but it's there if you know what to look for.
Oh wow, you're right, I just tried it out. But it's so blocky and apparently limited to (mostly) 90° geometry, that I'd never realized Google was doing anything but modeling a one-size-fits-all "rectangular corridor" along each street.
It seems like it's not exposing any kind of raw Lidar data, but a very simple geometric simplification of it. (E.g. trees are either ignored completely, or if there are enough of them they're treated as the side of a building instead.)
I completely understand what you're saying about the problems with smearing and warping due to not enough data. But I still can't help but wonder what it would look like if they were able to generate a "raw" (but denoised) Lidar geometry, so that trees and cars were treated as individual objects, rather than just either as part of the street floor or part of the building walls they way they are now.
* Need to do animation with very low latency (can't go to the network to collect data for the animation when the user clicks)
* Need to do the animation without much CPU/GPU power in the browser.
* Need to not download much data for the animation ahead of time (browsing panoramas is already very data heavy).