Moana Motunui Renderer on GPU
render-blog.com
render-blog.com
Chris are you on HN? I am curious whether some savings might be possible here by not snapshotting the IASes and GASes. Instead you can snapshot only the raw geometry, and then rebuild the GASes and IASes before every launch. It may be faster to rebuild the BVHs than to copy them from CPU ram!
*Oh incidentally looking through the code I also noticed that the GASes could be compacted which will save a lot of memory and time/bandwidth on the snapshot copies during rendering. Since you do compact the IAS, I am assuming it's problematic to try and compact the GASes?
In any case, nice work!
As for the lack of GAS compaction, I added IAS compaction to squeeze all of the beach's debris into a single snapshot, and just never got to GAS compaction. Right now each rendering pass is only computing one sample per pixel- so in addition to reducing the BVH footprints, I'm hoping I can also hide the transfer costs by computing many more samples per pass.
It’s been a while since I played with Moana’s meshes, so I don’t remember what GAS compaction will buy you exactly, but generally speaking it’s common to see the compacted GAS end up at about 50% of the size of the first build. If you’re allocating and building multiple GASes in parallel, then it might mean you have to defrag your snapshot in a separate pass using GAS relocation. Once you put all that together, it is possible that copying is faster than rebuilding, so just be aware I’m not certain that my suggestion to build every time is a winner in your case.
I can also envision some other strategies that may save considerably on snapshot copies in many cases ( and may be difficult to implement :) ). If you keep going with this and would like more support or suggestions or perf tips, please feel free to get in touch via the OptiX forum (Feel free to DM me there if you like). We’d be happy to try to help you squeeze out more samples per second.
This is open source: https://github.com/chellmuth/gpu-motunui/
Let me explain. GPU rendering of cinema quality scenes by itself is not that new. V-Ray GPU was around in 2013 and already had an impressive showreel back then: https://www.youtube.com/watch?v=RYPFY5OUzdk
I'd count Octane and Redshift as the next big contenders, who together with V-Ray switching from perpetual to subscription pricing, took over the market. Here's an example of a Redshift render from 2016: https://www.youtube.com/watch?v=to8yh83jlXg
Lately, Pixar has been handing out Renderman GPU betas, which has already been used in the new jungle book movie: https://youtu.be/tiWr5aqDeck?t=171
Along with this development, the price has gone down. From $1200 per V-Ray license, to $600 per Redshift license and now it seems like the previously prohibitively expensive Pixar renderer will soon be offered for $500.
And recently, people have been using Houdini ($6000 per seat) to export their scenes as USD (free open file format) so that they can use Blender 2.81 for rendering (also free). But the conversion to get things into Blender is work. This renderer can use the Disney production data directly, which is very convenient if you want to drop it in at the last minute for cost saving.
GPU rendering has gone from novelty in 2013 to established in 2016 to become the new default in 2018/2019. And this released as open source to me implies that the commercial renderer market will soon be dead. It looks "good enough" for advertisements and archviz, which is where the money is made.
That’s now plenty enough for all the geometry of a production scene plus a texture cache, but almost never enough for the base textures themselves.
I like the “use the cpu to do texture lookups while the gpu is tracing shadow rays” solution here. But I thought you or Keith had done a similar out-of-core with texture caching implementation for inclusion directly into OptiX. My assumption is that once you got the texture cache full enough, stalling on a Unified Memory page fault or similar wouldn’t be super common (again, depends on the hit rate).
Does the new cuMemAddressReserve combined with mmap on the host side let you have a giant virtual mmap’ed host side memory region?
We are shipping some (open source) out of core texture loading with OptiX now, lead by Mark Leone. Definitely yes, the idea is that after a few re-launches you’ll have all the mipmap texture tiles you need resident and you can render many samples without a stall.
To be honest I’m not sure whether we’re using mmap for this. (I’ve had my head buried in curves.) Mark is calling the high level concept “cooperative paging”, meaning the GPU and CPU work together to handle page faults and fill new load requests. The user can decide / define what the page fault behavior is, and whether & how they might want to fill the requests. They’re currently working on cache eviction, which will be a big leap when it arrives.
This basic technique could be used in Chris’ Motunui to fill geometry as well - you could use standin boxes for the instances and a ray hit would fault & load the internal geometry. So, obviously, that’s pretty complicated and not necessarily guaranteed to fit in memory, but when it worked it could save a ton of time. I’m thinking about ways to blend what Chris did with something like this to try and get the advantages of both.
Especially impressed as I've watched this movie at least 100 times over the last year. My 20 month old daughter is infatuated with it (not to worry, we usually limit viewing to a song-scene or two per sit down). The songs are involuntarily memorized and cries of "mana song" (sic) can be heard each time the toddler gets into the car. To the point that my wife, 3 month old daughter, the oldest, and myself are dressing up for Halloween as TeFiti, kakamora, Moana, and Maui, respectively.
If the OP should see this - what prompted you to want to attempt it, was it just the availability of the dataset or were you drawn into the story by screaming children as well?
Because the question on my mind is, given the improvements in rendering, there's some year in which we can now render in real time, what took a long time for Pixar to render.
So, what is it? I believe I read that creating the 3D version of Toy Story ran at 24 fps, averaged across the entire film. So, basically, we already crossed the threshold for Toy Story. (But maybe that was only if you have a rendering farm?)
So, where are we at? Could we render a plausible Toy Story in real time now? Bug's Life? Monsters Inc?
Last month they discovered a bug in Blender that was the cause of very slow hair rendering. The fix dropped render times for some scenes from 4 minutes to 40 seconds.
So it can indeed be useful to check if things can be improved for old movies.
But I am not sure realtime is an option right now. Most movies also use post processing per frame wich also takes some time.
The Kingdom Hearts games have areas based on various Disney IPs, and the last one came out fairly recently so it makes for a nice comparison
Pixar Toy Story vs Realtime Toy Story: https://www.youtube.com/watch?v=tkDadVrBr1Y
Disney Frozen vs Realtime Frozen: https://www.youtube.com/watch?v=q8vCQqg6SYg
Also, TS was definitely rendered higher than 1080p to make the transfer to 35mm film for large screens.
https://stack.com.au/film-tv/movie-review/toy-story-4k-ultra....
Since this is path tracing, how does this work in terms of instances from one batch influencing the lighting of a different batch?
- Matt Pharr (author of pbrt) -- 6 part series: https://pharr.org/matt/blog/2018/07/08/moana-island-pbrt-1.h...
- Joe Schutte (WDAS): https://schuttejoe.github.io/post/disneybsdf/
Their conversion script for the original Disney asset is available on GiLab: https://gitlab.com/3Delight/moana-to-nsi
https://pharr.org/matt/blog/2020/08/19/pbrt-v4-released.html
Is this part of the limitations the author discloses?: «Other features of the scene are possible to render but out of my initial scope, notably subdivision surfaces and their displacement maps, and a full Disney BSDF implementation»
In addition to dahart’s suggestions, I’d be curious to see some of your profiling that you did. Like, what’s the breakdown of the timing now? Are you PCIe limited, or have you overlapped the I/O with enough compute that it’s only XX% overhead?
How much memory do you need for all the geometry and BVHs? (Particularly post compression). I’m curious if this just barely fits on an A100 or similar.
Again, nice work!
https://arxiv.org/pdf/2001.02620.pdf:
"Although OSPRay already supported textures, it only supported image textures, however Ptex is a geometry based texture format baked on top of the underlying meshes [Burley and Lacewell 2008]. Not only does this mean there is no reasonable way these textures could be converted to 2D images for use in OSPRay, but that OSPRay’s entire view of how textures can be applied to geometry—which was inherently based on image textures—would have to change."
Total rendering time using a 12 core / 24 thread Google Compute Engine instance running at 2 GHz with the latest version of pbrt-v3 was 1h 44m 45s.
"When rendering the latest Moana island scene, pbrt-v3 uses 81 GB of RAM to store the scene description. Today’s pbrt-next uses 41 GB"
Matt's blog talks a _lot_ about all the optimizations he had to do just to get the scene to _load_ in PBRT. It's a beast.
Is it practical/useful? Let's put the timing in perspective.
A commercial CPU production renderer, 3Delight, has timing for the Moana asset rendered at 4k on their website.[1]
Time: ~34 minutes.
I asked them for details about the settings they used before posting this as the page only lists resolution.
4k resolution, 64 (shading) samples per pixel (spp), ray depths: diffuse 2, specular 2, refraction 4 (or 3, 3, 5, depending how you count ray depth). Machine was a contemporary 24 core server at the end of 2018. Mind you, the image is fine with 64 spp. Spp are hard to compare between renderers because optimizing path tracers is a lot about sampling. One renderer will converge to something useable with 1k samples while another just needs 64.
The 3Delight example is rendering all the geometry as subdivision surfaces with displacement (and their own Ptex implementation for texture lookups).
Timing comparisons of a different scene with recent 24core desktop AMD CPUs suggest that this asset would render much faster in 2020.[2]
The timing shows the issue with GPUs vs CPUs for this kind of assets. 5h for a 1k (!) resolution image with <= 5 bounces and 1024 spp (samples per pixel). That is terrible. Not using the real (subdivision) geometry and not using displacement.
I would love to see a breakdown how much of these 5hs is owed to the fact that the data doesn't fit on the device.
Using subdivision surfaces and displacement mapping make the amount of geometry grow exponentially. I.e. the out of core handling would predictably take an exponentially larger part of the render time.
Looking at the numbers I regularly get to see when counseling VFX companies on their rendering pipelines I don't see GPU offline rendering going anywhere for complex scenes. And even for simpler scenes where GPUs have an advantage still – with the CPUs AMD is putting out recently the gap is becoming very tight and if you do the math you often pay dearly for having an image a few minutes earlier (not even double digit minutes or hours earlier). Regardless of what Nvidia's marketing and some vendors who IMHO wasted years optimizing their renderers for a moving hardware target may want you to believe.
Regarding the latter: another point to consider is that you need to spend time working with/around the hardware limitations/bottlenecks of GPUs for this very the "scene doesn't fit on device" use case. Someone writing a CPU renderer can spend that time working on the actual renderer itself. This kind of software takes years to develop. Go figure.
Finally, as I expect this to be downvoted because of what I just said: don't take my word for any of the above. Just try it yourself.
The Moana Asset can be downloaded at [3]. A script to convert the entire asset and launch a 3Delight render can be had at [4]. The unlimited core version of the renderer can be downloaded for free, after registering with your email, at [5].
It renders with any number of cores your box has but it adds a watermark if no license is available.
Or you thy their cloud rendering. You get 1,000 free 24 core server minutes. Which is plenty to run this test.
[1] https://www.3delight.com/documentation/display/3DLC/Cloud+Re...
[2] https://www.3delight.com/page/features/2020-10-06-CPUbenchma...
[3] https://disneyanimation.com/resources/moana-island-scene/
The algorithm here to do out of core rendering is the important part, and it doesn't make sense for you to try to compare in-core CPU rendering to out-of-core GPU rendering.
> I would love to see a breakdown how much of these 5hs is owed to the fact that the data doesn't fit on the device.
I already know it's pretty close to 100% of the time spent handling out-of-core requests, that's not surprising, nor is it a bad thing (though it probably can be improved). If it were a CPU renderer doing this - streaming the geometry & BVH, the result would be the same (or much worse if streaming from SSD instead of an external ram of some sort.)