Vectorized Production Path Tracing
tabellion.org
tabellion.org
HPG Slides: http://www.highperformancegraphics.org/wp-content/uploads/20...
SIGGRAPH Course Notes (the relevant part begins on Page 35): https://jo.dreggn.org/path-tracing-in-production/part1.pdf
SIGGRAPH Slides: https://jo.dreggn.org/path-tracing-in-production/MoonrayV3.p...
Also, the actual shading system they used was also presented at SIGGRAPH 2017: http://blog.selfshadow.com/publications/s2017-shading-course...
They also mention the fact that programming the system is hard, and plugins must fall back on single-lane code paths until they can be coded into the system proper.
I would assume all the ray sorting makes it extremely difficult to use any bidirectional methods as well.
Also, path tracing may superficially at first glance look like a good fit for the GPU, but once you look closer, it becomes less so. With incoherent rays and surface shaders, memory access quickly becomes the bottleneck and the GPU cores stall. There are approaches to improving that (batching and sorting), but those can add quite some complexity.
Out-of-core memory for GPU rendering can allow for scene sizes to exceed VRAM, but then memory access becomes even slower. Again, with sorting and moving things to VRAM on demand there are paths to reduce the impact of it all, but those are again not trivial given the incoherent access pattern of simulating millions random walks.
[source: I am a developer working on the Cycles render engine (Blender) for a production company.]
Things get a lot more difficult when the scene becomes more complex, if it needs to be animated, etc though. You need much more advanced forms of visibility detection/object culling, track scene changes, need to have an much bigger working set (models, textures, metadata) in memory, etc. The amount of code that needs to run compared to 'just the path tracer' starts to far outweigh what the GPU can efficiently process.
Hybrid solutions are of possible and widespread, but it is not easy to implement those in a way that will not annihilate the speedup you get from the parts running on the GPU because of synchronization, copying memory around between CPU <-> GPU etc. I can imagine that the kinds of rendering pipelines animation studios use likely depend on hundreds of individual tools from different suppliers, so it would be next to impossible to integrate and the full rendering pipeline efficiently if it runs partly on CPU and partly on GPU. But maybe some parts could be?
It is used in some niches, and there are efforts to leverage both for interactive rendering scenarios.
edit: I think another reason might be something to do with the gigantic asset sizes (eg textures) typical in big studio jobs.