Raytracing on AMD’s RDNA 2/3, and Nvidia’s Turing and Pascal
chipsandcheese.com
chipsandcheese.com
https://forums.developer.nvidia.com/t/extracting-bvh-from-op...
As someone who deals with BVHs a lot for ray intersection, I find it pretty difficult to believe that leaf nodes with that number of primitives will be anywhere near performant, even with fast dedicated hardware like the RT cores.
It's true that the Nvidia cards have better intersection performance than ray/box tests, but I don't believe it's in the 100x ratio range which I suspect would be needed if the BVHs were that shallow and leaf nodes that large.
This is on Turing. Nvidia would've been motivated to de-risk the introduction of RTX by making boring choices. You may well see different results on later archs.
Previous indications from Nvidia about their BVHs don't seem to show anything about very shallow trees for any of the BVH algorithms that OptiX supports (scroll to bottom for reverse visualisation of a BVH hierarchy on top of the Stanford Bunny model): https://drive.google.com/file/d/1B5fNRFwv2LsGlCBJ8oKYRiiDUtL...
This really lays out the decisions that Nvidia made compared to AMD and how their approach tends to hide some of the shortcoming of GPUs (latency and utilization).
Lately their focus is ROCm and their CUDA equivalent language, but it also has limited official hardware support and AFAIK the Windows SDK for it is still not public.
Similar commitment issues have plagued their custom renderers.
Hopefully now that AMD has money they can turn things around. As much as it pains me to admit, giving up on OpenCL and focusing on the CUDA clone is probably the correct strategic move. OpenCL is just far too different to facilitate porting, and there's a lot of porting to do before AMD can start to take compute marketshare.
Now you have Nvidia CUDA, AMD "CUDA", Intel oneAPI. At least Intel is building their oneAPI on top of OpenCL via SYCL which means all AMD has to do is implement a bunch of Intel's OpenCL extensions, which are far from unreasonable (for example, unified shared memory which should have been in OpenCL 2.x instead of shared virtual memory).
BVH construction is my favorite question to ask in interviews because there's no single best solution and it mostly relies on mathy heuristics to get a decent tree. You can also always devote more time to making a more optimal tree but there's a tradeoff where it'll eventually take more time than it saves in raytracing.
I'd say RDNA 3 is not really giving useful ray tracing on for example 2560x1440 unless you use upscaling to speed it up. May be in a few GPU generations ray tracing will become usable with native resolutions.
When tray tracing will become fast enough without upscaling, then it will be a much better feature.
Personnaly, I am a dev, then I patch to compile out all that (and all the tracers at the same time) since ray tracing has currently a ridiculous ratio benefits/technical costs.
This defeats the very purpose of vulkan spirv: getting rid of those horrible high level shader compilers from the driver stack and keep them contained at the application level.
It seems beyond clumsy, but as I said, I need to get into the details of "why" those shaders in the first place, and then why they are not written directly in RDNA assembly or SPIR-V assembly (that would require an "assembler" coded in simple and plain C).