3D photographic inpainting from a single source image
shihmengli.github.io
shihmengli.github.io
I'm convinced that the folks that present at SIGGRAPH are capable of nothing short of black magic wizardry. I can understand the technology being used, but my visual cortex only sees the impossible becoming real.
This is truly deserving of the word awesome, because I am in awe.
You'll need to either use a separate 3D depth estimation AI, or (more likely) have someone do a manual stereoscopic 3D conversion of your historical image. Only then (when you have depth data) can the algorithm presented in this paper start its work.
They seem to be using MiDaS https://github.com/intel-isl/MiDaS for depth estimation which does a reasonable job on some random image pulled from pixabay https://i.imgur.com/IfbeaqY.jpg
Very cool application using Matterport I saw recently is to scan AirB&Bs, e.g. https://breconretreat.co.uk/accommodation/swn-y-nant/floor-p...
Except that there's 45 distinct planes and it directly interfaces with Unity, Unreal and ThreeJS in minutes.
Sorry if I'm reading too deeply into your note. There are just so many haters. Meanwhile, I backed these guys on Kickstarter, have had a unit on my desk for 18 months and think it's one of the most incredible things I've ever had to experiment with... and it cost me under $500.
I'm the author of a lenticular vray plugin and people love the technology for product design previews. Plus the fact that it's being mass-produced as children's toys makes it very affordable.
Are you asking for more information on https://lookingglassfactory.com/ or https://www.youtube.com/results?search_query=lenticular ?
That isn't what this is.
Watch the video in the bottom right corner, entitled "Comparison with State of the art".
Now go and rewatch the examples and actually look at the edges of the objects as they move the camera 'a bit'. You'll see a tonne of artifacting. Less, clearly, than the existing state of the art, so I tip my hat to the efforts here.
...but all this is doing is generating an 'empty gap' generated by perspective and then basically using the equivalent of photoshop's content aware fill to fill that gap with plausible pixels.
Since humans don't really pay much attention to edge details, it's quite plausible.
To quote the paper:
> In this work we present a new learning-based method that generates a 3D photo from an RGB-D input.
Ie. This work is taking a depth image as the input and working on that to generate a 3d photo, rebuilding the full content of each 2d layer in the image.
Ie. The output is not a 3d model, it is a series of 2D images at depth intervals, where the occluded content in each layer is in-painted (ie. generated artifically).
(NB. The 'from a single source image' work used here is not novel; they're just using existing approaches to estimate a depth image)
It might be "all" that it's doing, but it does it quite well and in a way that is quite believable, which is significantly better than what came before, which makes it almost realistic. That's what you said, but the way you said it felt like it lessened the achievement.
It's a bit better. That's the point I'm making; it's just incremental improvement on existing process. Read the actual paper, eg. under 'Quantitative comparison'.
If you think I'm belittling the effort I'm sorry, that's not my intention; ...but, for example, the other comments talking about using it to generate a full 3d model to display on a looking glass surface, or in VR displays a total lack of understanding of what has been achieved here.
So basically, if you have impossible to get input data, this network can do its magic.
What it then does is hallucinary inpainting, so something like the Photoshop content aware fill. If there's a tree in your photo, this one will make up a fake background behind it, so that you could move or remove the tree without things looking weird.
Except they evidently got the input data for the examples in the paper, so it can't be impossible to get.
They cite at least two different methods for adding depth information to a single image to generate the necessary RGBD data, different views of which can then be rendered with their inpainting applied:
As for using other AIs, they tend to not work too well on more complex images.
But in any case, getting the data that you need to be able to use the paper here is very challenging.
Edit: I should probably say that I have hands-on experience with MegaDepth and Midas and that it was underwhelming. Both of them assume a Gradient from close to far from the bottom to the top and both of them assume that optical variation will be in the foreground. A photo of a dining table from the side is already enough to confuse both of them.
But it is a problem well studied in machine learning and for which you can get off-the-shelf networks to use in conjunction with this paper.
Basically, it's similar to this but also re-lights reflective parts at the cost of needing more than one source image.
Just kidding. This is awesome! Just imagine the possibilities with enough computing power.
Maxed out at 4 GB RAM for 256x144 images for me.
PyTorch CPU installation instructions: https://pytorch.org/get-started/previous-versions/
OpenCV should work without CUDA. If not, build from source and consider `WITH_CUDA` flag.
That isn't to say they couldn't do it retroactively for single lens photos, but I'm guessing not right now they're not.
It's not like they come off as "let's remember the good ol' days with the separate water fountains for the colored folks." Not to me anyway.