27 karma · joined February 28, 2024
I haven't tried benchmarking SPIR-V shaders on macOS. Since they're translated into Metal shaders anyways, it should be possible theoretically.
Also, for the command line viewer in the new version, I've only tested make or ninja. I'll take a look at xcode when I get a chance.
Update: I just gave Instruments a try and it seems like the Metal compiler grouped all of the compute and copy operations together and just left the timestamp operations to run back to back. Since MoltenVK isn't a conformant implementation, I'm guessing the synchronization dependencies weren't respected.
However, I'm still getting 200ms frame times on the Garden scene at 4K with M1 Pro. The lego scene shouldn't be too bad even on an Intel Mackbook.
For my Apple Silicon benchmarks, the main bottleneck is the parallel radix sort that sorts the Gaussians by tile and depth. I used a some shaders from a sorting library, but it has some performance gaps with SOTA parallel sort algorithms. I think fixing this would give a 1.5x overall performance boost and maybe 3x on Macbooks. Also the wave size isn't tuned for different GPUs.
Another area of improvement is better management of the shared memory. Right now, we just let the driver manage it as the L1 cache. However, we could manage it manually and group Gaussian retrievals together for the same tile. This is what the official implementation does.
Although 3DGS is the first radiance field with SOTA quality that runs in real-time, I think it's still quite heavy. Due to the explicit representation of the scene, a lot of operations are memory bound. If you can't get an interactive frame rate right now, it's unlikely the improvements will make a material difference.
Hopefully that's where your work on compression comes in and solves the problem :)
Right now, the renderer runs on Windows, Linux, macOS, iOS, and visionOS (as an iPad app). OpenXR support and an immersive visionOS app are coming soon. Training is also WIP. As we're seeing the industry adopt research in Gaussian Splatting at a fast pace, I hope this makes it easier for folks implementing Gaussian Splatting or variants in their products.
Would love to hear your feedback!
For more context, see previous HN discussions on Gaussian Splatting: