Not at all an expert in this, but I'm curious how this compares to non-NN solutions like https://research.nvidia.com/publication/efficient-sparse-vox...
Presumably it takes less memory, letting a more complex scene be transferred to the GPU more quickly, at the tradeoff of time spent training the model?