Nvidia Research Turns 2D Photos into 3D Scenes
blogs.nvidia.com
blogs.nvidia.com
As I’ve gotten older, and my parents get older as well, I’ve been thinking more about what my life will be like in old age (and beyond too). I’ve also been thinking what I would want “heaven” to be. Eternal life doesn’t appeal to me much. Imagine living a quadrillion years. Even as a god, that would be miserable. That would be (by my rough estimate) the equivalent of 500 times the cumulative lifespans of all humans who have ever lived.
What I would really like is to see my parents and my beloved dog again, decades after they have past (along with any living ones at that time). Being able to see them and speak to them one last time at the end of my life before fading into eternal darkness would be how I would want to go.
Anyway, there’s a free startup idea for anyone—recreate loved ones in VR so people can see them again.
Always good to treasure the time we are given.
Let's be real though, the startup that makes this but appeals to our worst instincts make bank. I can't imagine how much more messed up future generations will be as we keep making more dangerous technology that appeals to our primal instincts.
It is a factual matter, easily discovered through simple searches, that non-Western cultures often take a different approach. Including, literally, ancestor worship. I am not making a judgement on which is better. I believe there are multiple healthy ways of dealing with grief, and "mourn & move on" can be one, but not the only one.
This is very far from an obsessive interaction with a semi-emulated version of a dead person in VR. However things wouldn't have to be that extreme.
In cultures that practice it, it's not uncommon for a small shrine in the home to be setup. And yet that is of limited access to family located further away, so a digital form of this not bound to a specific location could also be of use to people from these cultures. There is no reason that such things couldn't compliment existing practices of honoring ancestors.
You are absolutely right though that there are different ways of dealing with grief!
This seems very subjective, I don't agree at all.
That said, what would you do with your quadrillion year long lifespan, assuming you're healthy during it? It's way past everything you can do and learn, and cosmic level events are not what the human mind can perceive as they slowly unfold...
I would like to live a couple thousand years though, as long as my loved ones get to be along for the ride. Though I wonder if my loved ones would eventually become my hated ones...
I don't know whether this is something I would wish for myself. Maybe. I like being human. But maybe, who knows?
Works of art that require thousands of lifetimes to complete. Deep understanding of the physical world. Exploring the universe. Much more.
Re: creating new beings. How long till you also get bored of that? Remember we're talking about quadrillions of years.
An additional thing to consider: if you could live quadrillions of years and still be able to perceive things that take thousands or millions of years to unfold, wouldn't you subjectively be living the lifespan of a normal human being? That is, your life wouldn't feel long to your mind, but normal (your mind would have to adjust to run more slowly). Wouldn't then regret your "short" lifespan and wish you could live longer?
Such arrogance, humanity isn't all that far beyond it's banging rocks together phase. We haven't even scratched the surface. There is so much more to learn.
No one has ever had the experience to be able to say.
I would much MUCH rather live much longer. I don't hate my life at all.
Possibly where the Alzheimer's patients are told the same joke over and over, and it does bring them joy, comes to mind.
Remember, jokes about tinder/grinder dates didn’t even exist a decade ago.
Once a person dies, they're gone. The world is different, and nothing will bring them back. Spending time with a simulacrum isn't really spending time with a loved one.
I mean, already in recent years people have made some very low fidelity 'resurrections' and gotten some measure of comfort from it, never mind the many years of history of people who visit gravestones to 'chat' (some even believing they get replies to some extent or another). When markov chains were hip, "talk with Charles Dickens!" (play with a markov model trained on his works) was at least interesting to some, GPT of course can do a lot better. Imagine we had actual superintelligent AI working on this, which actually tried to recreate brain models, or even restore a sense of existence to the resurrection that can continue independently and to their own delight rather than just being some static VR experience turned on and off whenever. (I jotted some thoughts a few years ago... https://www.thejach.com/view/2017/6/how_you_might_see_some_o...)
I'm in agreement though at least for now that even if you dialed up the fidelity to the point I can't point to anything objectively "off", let alone went with a low fidelity VR thing or just a chat bot, I'd still always think in the back of my mind that this isn't really the same person. Nevertheless, it could still be comforting to some, and interesting to others, and so if it's at all possible for potential future humans/ems/aligned AIs to work at it, they will.
Curiously I don't have the same back of the mind feeling at all when considering the idea of someone preserved with cryonics and then brought back as an em, the person would be the same to me, even if there was a bit of damage from the vitrification process that was error-corrected. I have a disagreement with a friend on that point, who thinks it would be something similar but not the same. Anyway, I think this is due to just having a much higher fidelity "source of truth" to work from that is the preserved adult brain, whereas someone supposedly information-theoretically dead requires a lot more guess-work, perhaps truly impossibly too much more even for a superintelligence, to bring back convincingly.
For less convincing avatars where the point is just my own benefit of conversing with something like them when I want, from slightly like them to eerily like them, for one it's a weak yes, for the others I'm more indifferent -- it'd be more in the realm of curiosity than desire, like talking to a historical figure or a fictional character. The weak yes I expect will get weaker (as it already has, despite non-linear flare-ups/resurgences where it's temporarily stronger) and eventually match the others after long enough.
Will the simulacrum age? Will it change? Will it ever surprise me or intrigue me? And if it does, is it something the dead person truly would have done?
It's sort of like a photograph; a photograph seems to capture reality but all it shows is some abstraction of a physical reality at one point in time. The photograph tells me nothing about the current reality of anything depicted within.
A simulacrum gives me the person I knew, when they died, and that's it. Perhaps comforting and interesting, but ultimately unfulfilling and unsettling I would imagine.
Maybe that's fine for some people. To me there's a line where it crosses over into offense. In no way do I think the status quo of my mental well-being is so important that I'd replace someone with a digital robot facsimile.
"You know, you should really do XYZ..." (whatever the agenda is)
Pretty dystopian.:(
Relevant Asimov story: https://en.wikipedia.org/wiki/The_Dead_Past
Reading about The Dead Past tho, there's a good deal of overlap with the plot device of "Devs", a periscope into deep history, but which Nick Offerman's character uses to re-live moments of his daughter's short life, as Asimov's protagonist does.
I for one have many things I wish I had said to my parents before they died. I’d like to use it, but not be obsessed or consumed by it.
please tell me more!
Some interesting replies there
I don't even know if it would have some kind of therapeutic use of whatever. Yeah ir would be a wonderful experience and I... ehh... let's say "dreamed" about similar experiences a few times and it was a powerful "dream". But being able to do it on demand would I think change how grieving works in the modern world so much! And we rely on the things that went to not actually be there for us on demand to be able to move on.
It's a beautiful idea but it's as beautiful as the fear of nuclear MAD keeping the first world safe from war in their own land, not as beautiful as a poem or a flower or a good memory.
That isn't necessarily what eternal life would be. Indeed, eternity is not temporal at all.
For example, in the Catholic understanding of heaven, those in heaven exist in aeveternity or aevum[0], a state in between the temporal and the eternal. Eternity[1] is proper only to God and is by definition timeless, with no beginning or end, so it would not make sense to speak of quadrillions of years. Where God is concerned, there is, loosely, only a now with no beginning, no end, no past, and no future, only the present, to use temporal language analogically.
Furthermore, heaven is in part characterized by the beatific vision[2] which is an immediate, direct, and inexhaustible knowing of God (Being Himself) which is Man's ultimate and supreme happiness. In this life, knowledge of God is generally mediate like much of human knowledge.
In other words, thus understood, the best of this life is but a faint shadow of a shadow in comparison to the ultimate fulfillment of heaven which is Man's proper end.
[0] https://www.newadvent.org/summa/1010.htm#article5
That’s probably the basis for most of human religion and philosophy; coping mechanisms.
I’ve always wondered if it might be possible to ”reverse engineer” parents or siblings from one’s own DNA.. at least visually.
[1] https://www.sfchronicle.com/projects/2021/jessica-simulation...
“Yes what a wonderful day. I also fondly remember the Adidas sneakers I wore. Did you know they’re 25% off at Target this weekend only?”
Tim Heidecker did this with his son Tom Cruise, Jr. Heidecker and it was quite moving: https://youtu.be/2fVQtmMN_ao?t=178
- Shooting a movie from a few cameras, creating a movie version of a NeRF using those angles, and then dynamically adding in other shots in post
- Using lighting and depth information embedded in NeRFs to assist in lighting/integrating CG elements
- Using NeRFs to generate virtual sets on LED walls (like those on The Mandalorian) from just a couple of photos of a location or a couple of renders of a scene (currently, the sets have to be built in a game engine and optimized for real time performance).
Having said that, it might be the end for any junior type of roles. Same reason that github copilot really takes a bite of the need to have a junior developer.
I'm very curious what will happen because it will become a sort of trend across other industries apart from legal or medical professions (peace of mind from human-in-the-loop).
The real trick is textures.
When the photo is laid over and wrapped on the model, it looks great. But once you remove that, the raw mesh underneath is not as impressive.
I would really like to see the examples they had here without the texture laid on top.
The cool part is that this also allows for capturing transparency, and any effects caused by lighting (including complex specular reflections) are embedded into the representation.
There are examples of people outputting SDF (and by extension geometry) with nerf, and projecting original texture onto that would give some nice effects; (live volumetric works best this way) though there would be some disparity where edges/occlusion isnt perfect, so youd want to sample nerf's rgb anyway... although a lot of that is fuzzy at the edges too. A lot of incorrect transparency at edges looks great in the 2D renders (so much anti aliasing and noise!) but less good for texturing
And while it's true that there are methods for extracting a surface from a NeRF, achieving a high quality result can be challenging because you have to figure out what to do with regions that have low occupancy (i.e. regions that are translucent). Should you consider those regions as contained within the surface, or outside of it? Especially when dealing with things like hair, it's not obvious how to construct a surface based on a NeRF.
Having 3d renders of the entire film without needing green screens and a bunch of balls seems like it would have to make some of the post processing work easier. You can add or remove elements. Adjust the camera angles. More effectively de-age actors. Heck, even create scenes whole cloth if an actor unexpectedly dies (since you still have their model).
Seems like you could also save some time having fewer takes. What you can fix in post would be dramatically expanded.
Best part for film makers, they are often using multiple cameras anyways. So this doesn't seem like it'd be too much of a stretch.
> - Using NeRFs to generate virtual sets on LED walls
Sounds like a powerful set of tools to defeat a number of image manipulation detection tricks, with limited effort once the process is set up as routine. State actor level information warfare will soon be a class of its own. Not just in terms of getting harder to detect, but more importantly in terms of becoming able to produce "quality" in high volume.
I wonder what happens to most people when they see innovation such as this. Over the years I have seen numerous mind-blowing AI achievement, which essentially feel like miracles. Yet literally after an hour I forget what I even saw. I don't find these innovations to have a lasting impression on me or on the internet except for the times when these solutions are released to the public for tinkering and they end up failing catastrophically.
I remember having the same feeling about chatbots and TTS technology literally ages ago, but at present time, the practical use of these innovation feel very mediocre.
> Big news! I sent @sdw the original image. He theorized a leaf from a foreground tree obscured the face. I didn’t think anything was in view, and the closest tree is a Japanese Maple (smaller leaves). But he’s right! Here’s a video I just shot, showing the parallax. Wow!
TIL though, thank you!
For 3D modelers, this is huge since it takes a lot of experience and grunt work to put the right touches to get an even a boilerplate 3D model. So much so that many game companies have outsourced non-human 3d modeling, this would certainly impact those markets.
1) It could further lower the cost and improve quality.
2) Studios could move back those time-consuming tasks on-shore and put an experienced in house artist/modeler to manage the production.
3) Hybrid of both
What I see here is that NeRF has a far more impact to the 3d modeling/animating industry than github copilot. Another certainty is that we are going to see faster rate innovation. We are at a point where a paper released merely months ago are being completely outpaced by another. The improvement in training time that NeRF offers is insane, especially given how quickly this new approach came out.
We could be at a future where the release of AI achievements will not be able to keep up with published works. It would be as fast as somebody tweeting a new technique, only to be outdone by somebody weeks or possibly days.
Truly exciting times.
(not to say this sort of behavior is exclusive to corporate PR. as the best and smartest person ever, I would never need to exaggerate my achievements on a job application, but others may)
But then once it IS released to the general public, it's probably been at least several months, maybe even multiple years since the announcement, so people are like, "yawn, this is old news."
Another comment mentioned education: I take classes, formal education, for things that are immediately applicable for me right at the time
So its pretty consistent for me
Another thing that might be different is that I don't go for perfect or good enough, or worry that an existing alternative like "having an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" might already exist, I just go for novelty and am selling to people that also like novelty
Neural TTS is so good now it's being used to fake the original voice actors in video game mods.
My impression is that you take a bunch of photos in various places and directions, then you use those as samples of a 3D function that describes the full scene, and optimize a neural network to minimize the difference between the true light field and what's described by the network. An approximation of the actual function, that fits the training data. The millions of coefficients are seen as a black box that somehow describes the scene when combined in a certain way, I guess mapping a camera pose to a rendered image? But why would that be better than some other data structure, like a mesh, a point cloud, or signed distance field, where you have the scene as structured data you can reason about? What happens if you want to animate part of a NeRF, or crop it, or change it in any way? Do you have to throw away all trained coefficients and start again from training data?
Can you use this method as a part of a more traditional photogrammetry pipeline and extract the result as a regular mesh? Nvidia seems to suggest that NeRFs are in some way better than meshes, but according to my flawed understanding they just seem unwieldy.
You don’t change NeRF (the model). You change the point of view of an observer.
With other structured data compositing and animating is relatively trivial.
It turns out that people have approached this problem before and you can composite nerf too (1) by sampling different functions over the volume.
…but, let’s not pretend.
The complaint is entirely valid. You’re taking a high resolution voxel grid and encoding it into a model.
Working with simple voxel data let’s you do all kinds of normal image manipulation techniques, and it’s not clear how you would do some of those with a nerf.
Practically speaking, the applications you can use this for are therefore reasonably limited right now.
[1] - https://www.unite.ai/st-nerf-compositing-and-editing-for-vid...
Any image transformation you can do on voxels you can straightforwardly transfer to nerfs. Voxel data is just a lookup table from discrete positions to material properties like color and density. When you apply a transform, you change the inputs (e.g. multiplying them with a rotation matrix) or the outputs (e.g. changing the color). If you want to do the same thing with a nerf that maps continuous positions and directions to material properties like color and density, just transform the inputs or the outputs.
The major difference is that with voxel data you can easily do output-modifying transformations directly on the stored representation, while for nerfs it might be cheaper to do it on the fly instead of redoing the training procedure to bake the change into the model.
No.
> it might be cheaper to do it on the fly instead of redoing the training procedure to bake the change into the model.
I think it’s a bit more complex than you imagine; it’s not “cheaper/not cheaper”; it’s literally the only way of doing it.
If you have a transformation f(x) that takes a pixel array as in input and returns a pixel array as an output, that is a trivial transformation.
If you have a transformation that takes a vector input f(x) and returns a pixel output, it’s seriously bloody hard to convert it to a “good” vector again.
Consider taking a layered svg and applying a box blur.
Now you want an svg again.
It’s not a trivial problem. Lines blur and merge, you have reconstruct an entirely new svg.
Now you add the constraint in 3d; you can never have a full voxel representation in memory even temporarily because of memory constraints.
At best you’re looking at applying voxel level transformations on the fly to render specific views, and then retrain those into a new nerf model.
I think that counts as … not straightforward.
Doing all your transformations on the fly is a lovely idea, but you gotta understand reason nerf exists is that the raw voxel data is too big to store in memory. It’s simply not possible you can dynamically run a image processing pipeline over that volume data in real-time. You have to bake it into nerfs to use it at all.
> Now you want an svg again.
> It’s not a trivial problem. Lines blur and merge, you have reconstruct an entirely new svg.
Nerfs are not svgs.
Consider taking a nerf and applying a box blur. Easy, a box blur on voxel data takes multiple samples within a box and averages them together, so to do the same thing to a nerf just take multiple samples and average them together.
That does get slower the more samples you need, but you never have to materialize a full voxel representation.
> Doing all your transformations on the fly is a lovely idea, but you gotta understand reason nerf exists is that the raw voxel data is too big to store in memory.
> It’s simply not possible you can dynamically run a image processing pipeline over that volume data in real-time.
Photogrammetry is great if you have a very solid object that is not shiny or translucent at all. You get a lot of surface color micro-detail and a bit of bumpy meso-detail.
But, if something is fuzzy, hairy, or lacey or smokey you are straight-up out of luck. Don't even try.
If it is shiny, it can be difficult to capture at all --let alone capture the shine. Material capture techniques that are not "chalk sculpture" are rare, very limited and usually experimental.
NeRFs however are pretty much a photograph that you can walk around in. They have about as much structure as a photograph ;) But, that lets them not care about the mathematical definition of your scuffed-up, lacquered, iridescent, carbon fiber mirror frame and just show it as it looks from whatever angle.
My main problem with them is that it seems as if all the data is unstructured and interdependent, not like pixels, voxels, or similar where you can clearly extract and manipulate parts of the data and know what it means. To use your photograph example, a digital photo is a simple grid of colored points, and it's easy to change them individually. A regular 3D scene is a collection of well defined vertices, triangles, materials etc, that is then rendered into a digital photo using a easy to describe process. A NeRF on the other hand seems to be more like enter camera pose => magic => inferred image.
Maybe I'm overthinking it and it doesn't have to be as general as out current formats, maybe a binary blob that can represent a static scene is fine for plenty of applications. But it feels needlessly complicated.
You can generate a mesh from the density information this function gives you, but for that you need to discretise the continuous densities you get out.
In my opinion, Nerf is more about showing progress in making AI memorize 3D scenes and the hope is that this will lead to actual understanding sometime in the future.
The research is moving fast though, so if you want something almost as fast without specialized CUDA kernels (just plain pytorch) you're in luck: https://github.com/apchenstu/TensoRF
As a bonus you also get a more compact representation of the scene.
Generating the novel viewpoints is almost fast enough for VR, assuming you're tethered to a desktop computer with whatever GPUs they're using (probably the best setup possible).
The holy grail (from my estimation) is getting both the training and the rendering to fit into a VR frame budget. They'll probably achieve it soon with some very clever techniques that only require differential re-training as the scene changes. The result will be a VR experience with live people and objects that feels photorealistic, because it essentially is based on real photos.
Doesn't seem like much of a stretch to determine the angles as well.
E.g. a semi brute forced way with GANs
It all works pretty well. Trying it on your own video is pretty straightforward.
If you're genuinely curious, look into structure from motion, visual odometry, or SLAM.
Not really, with SLAM there are various algorithms to keep inaccuracy in check. Basically it works by a feedback loop of guessing an estimate for position and then updating it using landmarks.
If you want to experiment, take a bunch (~100) of photos of an object, and use COLMAP to generate the poses. COLMAP implements a global SfM technique, so it will be very accurate but very slow.
But NERFs can capture and render things that are very difficult subjects for photogrammetry: Fur, vegetation, reflective or transparent surfaces etc.
https://github.com/yenchenlin/awesome-NeRF
It's possible to capture video / movement to into NeRFs, possible to animate, relight, compose multiple NeRF scenes, and a lot of papers are about making faster more efficient and higher quality NeRF. Looks very promising.
NERF is a very active research area, and the progress from 2020 to now has been nothing sort of astonishing. In 5 years, I expect there to be fully generative NERF's in research i.e describe a scene, and a NN produces a full 3d scene that can you interact with.
One thing I noticed is there are no pedestrians and cars in those scenes. So they must do a lot of work to filter them out by combining a lot of footage. Therefore, it likely can't be used (as-is) on the street view dataset...
Certainly though, game franchised films would become a lot more imersive, though I do hope that whole avenue dosn't become sameish with this tech overly learned upon.
But one thing for sure, I can't wait to bullet-time the film - The Wizzard of OZ with this tech :).
The days of Meta being a giant that can continously buyout other companies to keep stayling alive are coming to an end.
Why they don't explain the scope of achievement properly?
edit: I don't think it is just 4 https://news.ycombinator.com/item?id=30810885
The model requires just seconds to train on a few dozen still photos — plus data on the camera angles they were taken from — and can then render the resulting 3D scene within tens of milliseconds
Pretty impressive, but lesser compared to generating it from 4 photos (which imho the movie suggests). Which would be "real magic level of impressiveness" for meHow much it resources it takes to generate images like that? is this the most ideal situation?
Can you take images from the web and based on metadata make a better street view?
With all this AI where is one accessible translation service? or even an accent-adjusting service? or just good auto-subtitles?
This just isn't true. I can create a 3D scene from 360-degree photos (even 4) in a minute or so using traditional methods, even open-source toolkits.
It doesn't look as good as this because it doesn't have a neural net smoothing the gaps, but it's not true that it takes hours to build 3D information from 2D images.
I highly recommend trying this at home:
https://nvlabs.github.io/instant-ngp/
https://github.com/NVlabs/instant-ngp
Very straightforward and gives better insight into what NeRF is than any shiny marketing demo.
To that end, I wonder if exporting a similar bit of video from that same path exported as stills would be enough to generate the 3D version.
Now that I had a 3d model of the scene, I had to spend countless hours cleaning it up to make sure it was useable. Maybe in the last 5 years, things have improved.
But this demo used 4 pictures. And apparently, it rendered the final image in seconds. That's what's new.
Was that created from a few photos? I didn't see any additional imagery below
--- Update
It looks like these are the four source photos: https://blogs.nvidia.com/wp-content/uploads/2022/03/NVIDIA-R...
Then it creates this 360 video from them: https://blogs.nvidia.com/wp-content/uploads/2022/03/2141864_...
four source photos
Is it just 4 or are there more?I find it hard to believe there is only 4. There are clearly more data in video
https://i2.paste.pics/645fe17e418b2cb1f6179e0b6671a170.png like back side of camera here (it is kinda visible but much poorer compared to video). Or existence of a 2nd white sheet in background. But correct me if it is only 4 and you have a source on that
I wonder if they could reach a trillion market cap.
Nah, if it is a joke at their own expense then it is "self deprecating humor", something which is definitely designed to get a laugh. Humiliation fetish, maybe? Obviously nothing is funny past a certain point of deconstruction... especially if you find yourself defending the distinguishing difference of the "meta". Just stop making puns, easy.
:P
They do high quality research and almost inevitably end up releasing the code and models. It's possible to reconstruct all that as a non-CUDA model, but when you want to use it, why would you when it's going to take months of work to get something that isn't as optimised?
Was talking to someone 2 days ago, just died randomly, early 40's. It's trippy, I have data of this person's face eg. videos/base64 strings... it's eerie. Unanswered texts wondering what's wrong. My thinking is I was only exposed to a part of this person, won't be them fully if reproduced.
If you detect an apple in a photo, you could quite reliably guess how the back look
Still very cool :)
I've been trying to get GANs to do this for a while, but NeRFs look like the perfect fit.
If it's that tech, I'm pretty sure it got dropped when 3D acceleration made it feasible to just re-render the whole scene every frame. Dedicated hardware won out over software tricks.
There was an article here some time back that showed streets in NYC that used this kind of idea of using older photographs to put one in street view of older NYC. So, yeah, I'm guessing it could be done. Might be weird with the different quality of images (modern digital, polaroids, kodachrome, etc).
I know this post adds nothing, but that one's well worth being pointed out.
(They could be stands for lights, but like I said, just guessing.)
I mean, obviously generated images can't be used as proof in the court of law, but this feels like we're slipping into crummy USA show territory.