PlenOctrees For Real-time Rendering of Neural Radiance Fields (NeRFs)
alexyu.net
alexyu.net
http://alexyu.net/plenoctrees/demo/?load=https://storage.goo...
https://phog.github.io/snerg/viewer/index.html?dir=https://s...
This matters largely because there is more to NeRFs than just the first paper; eg. there are Mip-NeRF (https://jonbarron.info/mipnerf/) and D-NeRF (https://nerfies.github.io/). A quantized bunch of samples loses the power of neural implicits.
See an overview of NERF here:
Do they produce geometry/polygon data as an output of the neural network?
This paper introduces a way to render NeRFs at reasonable speed. Still not stellar, but quick enough to make NeRFs useful for many use cases.
A NeRF takes over both the role of the file format and part of the rendering in the form of a neural network. You feed in a world coordinate and a viewpoint and you get an RGB tuple and density out of it. If you interrogate the NeRF enough you can render any traditional 2D or 3D image out of it by combining all the datapoints.
One theoretical benefit is that that a NeRF is a continuous function, so the resolution is only limited by the capacity of the neural network. Another cool thing is that a NeRF is trained on pictures (with info about where they were taken from), so if you train a NeRF successfully in high-res it’s like scanning an object. A major practical challenge is that it is (was?) pretty frickin’ slow to work with. I wrote a more elaborate comment about it on the previous NeRF improvement post [1]. There I closed with:
> It would be amazing to have NeRF-based graphics engines that can make up spaces out of layers of NeRFs, all probed in real-time.
Here they’ve taken a major step in that direction by speeding up the rendering 3000X.
It bakes the NeRF back to a semi-discrete representation (Octree of Spherical Harmonics voxels) which can render near-identical results at interactive speeds.
The baked data is much larger than the original NeRF model (2gb vs 5mb), but they can be downsampled to 30-100mb with little loss in quality.
This preserves the exact lighting equation that the NeRF learned, while traditional voxel rendering is limited to traditional lighting equations.
You would have a hard time voxelizing a NeRF, because you can't extract a traditional lighting equation out of it.
And in the paper: https://arxiv.org/pdf/2103.14024.pdf
> As training time poses another hurdle for adopting NeRFs in practice (taking 1-2 days to fully converge), we also showed that our PlenOctrees can accelerate effective training time for our NeRF-SH.
I have tried meshroom too, the result is equally bad.
As someone who consumes plenty of science fiction, I could easily imagine a sci-fi story that breaks information down into some epochs...
epoch 1, "analog": the incoming signal is the raw representation of the data, like grooves on a vinyl record or AM/FM radio
epoch 2, "digital": data is encoded in binary format, and possibly compressed
epoch 3 "???": you don't store the data itself, you store a neural network backed function that can reproduce the data in many different ways based on inputs. I.e. you don't store a video of a chair, you store a neural-net that can approximate the chair from any angle with any light source, PLUS you also store a "camera track" that can be used as a pre-defined input to get a preset "video". But at any time, the user can "unlink" the track and operate the "camera" as they choose.