Neural Supersampling for Real-Time Rendering
research.fb.com
research.fb.com
However Nvidia are treating DLSS as their secret sauce and not publishing any details, so Facebook's more open research is interesting even if it's not as refined yet.
From the inputs you can glean that DLSS is using temporal integration but that's hardly a new idea, the novel part is in how it performs the integration and Nvidia isn't sharing that.
You can look at the tradeoffs of neighborhood clamping/clipping here:
https://de45xmedrsdbp.cloudfront.net/Resources/files/Tempora...
Those slides were before TAA upsampling was added, but the issues are similar once it is in place too. Unreal Engine already has TAA upsampling, but has to fight against the issues mentioned in the Nvidia presentation.
Nvidia trained it for specific upsampling multiples, so the main functional difference with existing TAA upsampling is it can't do continuous base-resolution changes.
If he says that the GP is inaccurate, trust him.
Hi, Brian!
The part of the hardware that is running the DLSS ML model are the tensor cores. But the algorithm and the model itself is not baked in, it is provided in the driver and/or game
A Titan V GPU, using the 4x4 upsampling, at a target resolution of 1080p takes 24.42ms or 18.25ms for "fast" mode. This blows out the 11ms budget you have to render at 90hz (6.9ms for 144hz), and it doesn't appear to include rendering costs at all...that time is purely in upsampling.
Cool tech but a ways to go in order to make it useful for VR.
That part wouldn't be an issue if the plan is to render low resolution images in the cloud and stream them to a device that can upsample them locally. There wouldn't be any local rendering costs.
Any compression artifacts are going to stick out like a sore thumb, so you'll need to stream very high quality, and you're going to have weird interactions between different layers being compressed differently.
Spending precious milliseconds perfecting the corners of the image for VR seems like a complete waste.
FVR needs a hook: what can it do that "dumb" VR headsets don't?
the bit of your eye that need high resolution is surprisingly small.
ATW means that your head rotational latency is fixed at your display rate. Now the content still has a higher motion to photon delay. The user's perception of motion sickness though is not about that though. It appears to be primarily incongruity between where your eyes are focusing & the signal from your inner ear. ATW is relatively cheap & well understood so high refresh rates with a deep pipeline aren't totally unreasonable.
Once you have that, the total end-to-end latency matters in a different way. Now you're focused on the content not feeling laggy. That's where prediction comes in. 100ms of end-to-end latency is probably outside the realm of what is doable today. With algorithmic improvements, increases in computational power, improvements in body modelling & neural networks, it's not unreasonable to start thinking about 100ms.
Anyway, the mental picture I meant to paint you was what if you thought about all the content in terms of how mobile GPUs (& desktop GPUs under the hood) think about it, in terms of tiles. What if your physics simulation only bothered to simulate the current tile's worth of content before handing off to the rendering stage. What if rendering focused on that tile before moving onto the next? Then you submit a tile's worth of content to the display. Then your display isn't refreshing the entire panel uniformly but instead refreshing each tile offset from the other. On top of that you bake time warp directly into the display driver.
This is a monumental undertaking. The reason we haven't historically done that as an industry is because HW speeds with the existing architecture have increased far faster than the SW architecture could mature to take advantage, so from a cost perspective both the HW & SW vendors would be idiots to try to do this. However, HW speeds are petering out & even purpose-built accelerators can only do so much. At some point the architectural shift will happen.
I'm not trying to make this seem like a trivial thing. The single global view is easy to reason about. We need fundamental research into this space on how to build physics & rendering engines that can unlock this kind of granularity. There's massive data dependencies that are really challenging to untangle. Each layer in the pipe, from physics engines & rendering engines, to display pipeline APIs, to GPUs likes to treat everything else like an opaque system with each providing a complete result to the next. There's an enormous amount of existing architecture that would have to be rewritten. I'm also 100% clear that it's entirely possible that parts of the pipeline can't be segmented in this way to get good looking result or it might make the SW burden on content creators to costly. I still believe it's the right direction though & is the only way you're going to ever see 8k displays @ 120hz or above. Well, that's not strictly true. The other trick will be to use motion estimation & neural nets to provide the interpolation when you drive it higher than the content can support - you can see that with all the motion estimation & supersampling improvements that various vendors are producing.
I was instead saying that if you could re-architecture the entire system end-to-end so that you could generate 1/10th of the frame in 1/10th the time, & there wasn't a resource contention between frame (i.e. you would start processing the next frame as soon as you're finished with the current one), then even though it took 100ms for the entire frame to scan out, each individual section is still updating at 100 fps. Similarly, if you could get update the display in pieces rather than the one-shot serial operation as it is now, then you'd see some serious changes to our understanding of latency & pipelining.
This is all easier said than done & was a total strawman taken to an absolute extreme. However, if the SW could get that pipelined, HW vendors would have a much easier time making their end of it more efficient.
I suppose it's assumed that with the contributions of this one, future work can be done to make it faster.
> and combines the additional auxiliary information
> multiple frames
In other words, the label "Low Resolution Input" on the blurry images is misleading. The image should be labelled "some of the input".It uses color, depth and subpixel motion vectors of 1-4 previous frames. All things that modern game engines can easily calculate. You didn't even need to read the paper to get this info, it's literally in a picture on the blog post.
The low-res image is itself output generated from a lot of other input the game engine generates. That same input that is already being generated anyway can also be fed into this to improve the post-processing. Finding ways to productively reuse already existing/generated data is the hallmark of any top (graphically) game.
I read a detailed write-up on the graphics pipeline for GTA:V on Xbox 360. It blew my mind how many different ways they reused every single bit that ever hit the RAM. Which explains how they pulled off those graphics on a system with half as much RAM as an Apple Watch.
The inputs are similar:
https://www.nvidia.com/content/dam/en-zz/Solutions/geforce/n...
In contrast to DLSS1, the output of the NN is not color values, but sampling locations and weights, to look up the color values from the previous low-resolution frames.
Great start but definitely needs additional work to be usable in games.
The H.264 encoder on my CPU introduces >16.7ms of latency into a video stream, but it can encode hundreds of frames per second of SD video all day. Adding ~1 more frame of latency may be worth a quadrupling in image quality/resolution in most circumstances.