The problem is "how do you do 3D deep learning 3D scene reconstruction" aka "how to make 3d equivalent of stable diffusion".
Voxels are bad as they take too much memory and are uniform (each voxel is eg 5x5x5 cm).
Gausians are kinda like "variable sized voxels" that also have orientation, color based on angle viewing and can stretch.
Imagine 3D blobs basicaly (3D capsules or more like 3D density bubbles with transparency).
So the scene can be represented in 3D much more efficiently using gausians splats, which is why they are used for "3d stable diffusion".
So now that this has come into use we also need efficient way or rendering them.