3D Gaussian Splatting as Markov Chain Monte Carlo
ubc-vision.github.io
ubc-vision.github.io
> We initialize our samples either randomly or from point clouds, typically from Structure-from-Motion (SfM) as in 3DGS
If you have a rigid multi-camera rig, the camera poses might be known from calibration, but then the particular scene shot on such a rig, could be reconstructed without COLMAP or other structure-from-motion tools, if I understand it correctly.
Gaussian Splats are able to handle more heterogeneous information sources, allowing more sources to help splat an environment. Devices like drones, surveillance cameras, or autonomous systems can be used to create or incrementally update a Gaussian Splat; and there's interesting work to allow them to locate themselves within the splat, not just to show themselves but also to place vision ML outputs into it (such as object detection or segmentation results).
Up till now nearly all digital representations of physical environments are either based off the original designs (by things like CAD or BIM files), or are an approximation of the environment (from photogrammetry or Lidar scans). CAD and BIM files suffer from drift, the real environment almost never perfectly matches the design files, small (and large) changes are made; and many times those files aren't even available if the structure isn't new. Photogrammetry and Lidar scans struggle because their output is a pointcloud, and it's very difficult to accurately mesh a pointcloud (Matterport only partially solved this problem and sold for $1.6B). Gaussian Splats overcome these issues; they're comparatively easy to generate for any environment, and allow for very accurate and easy viewing from any angle.
I think the Digital Twin space will be turned upside down, and they could potentially even cause huge changes in autonomous and semi-autonomous factories, warehouses, and depots. A single Gaussian Splat could be the source of truth that many autonomous vehicles update through their separate SLAM systems. Operators then would have access to this splat (and it's history) as a source of truth for the environment. Then, using techniques like iComMa[1], it may be possible to directly align XR devices into the Gaussian Splat; allowing operators direct access to location-based information generated by the environment.
That's a lot of words to say: Gaussian Splatting is a very neat new technology that could really underpin many future technologies, I'm really excited about it
On the other hand, I'm playing with the idea of a platform which provides a Gaussian Splat based Digital Twin of an environment so other systems can utilize it to share location-based information. Even though I don't think it'll be possible to build without utilizing Gaussian Splatting; splatting may not end up in any of the pitches or advertising directly.
Splatting is fundamentally about viewing pointcloud data. That's great. But it doesn't deal with all the other functions virtual twins need pointclouds for (e.g. design vs real world conformance).
Pointclouds themselves are proving hugely useful in a number of fields but vary considerably in form and application often based on how they are captured (e.g. LiDAR vs SfM photogrammetry)
Visualising pointclouds effectively has been a major problem which splatting really solves elegantly so it will be a major practical advance when splatting is added to cad software and javascript map visualisation libraries.
what exactly do you mean by 'digital twin'? do you mean any kind of computer model of a real-world phenomenon, including sets of differential equations, as i've sometimes seen it used? presumably you mean something narrower, but how narrow? do you mean, for example, specifically cad models of parts that are going to be manufactured?
i guess this sounds like i'm nitpicking but actually i just want to know the scope of the space that you expect to be turned upside down
They have also included spherical harmonics for view dependence since day one.
There is nothing that prevents gaussian splatting from being used dynamically. There are a variety of approaches to extend gaussian splats into the time dimension to capture and represent a 3d scene over time. The challenges here about how to capture sufficient scene data (or use ai to fill in insufficient data) and how to compress it. There are also techniques that enable dynamic simulations, or real time animation of collections of splats.
Also, adding un-baked lighting to gaussian splats is not particularly hard, you can already throw slats into several game engines / 3d renderers and add new lights to them. The hard part of relighting is taking an existing capture of a scene with baked-in lighting and deriving the resulting material properties and lighting sources. This isn't directly related to gaussian splats themselves though, you would have a similar problem recovering the base materials and lights from a 3d mesh with baked-in lighting textures. This really falls under a separate category of techniques called "inverse rendering". If anything, gaussian splats give us a new tool to help with these sorts of problems.
Honestly the biggest remaining roadblocks to more elaborate and widespread uses of gaussians as a rendering method are probably storage and performance related. And I'm optimistic these will be convincingly solved, triangle rasterization has had many orders of magnitude more research, optimization, and custom hardware built around it.
https://computergraphics.stackexchange.com/questions/4164/wh... is an introduction
Is there something I can download on my Meta Quest to try it out?
But also you're just throwing a ton of primitives at the problem; high quality scenes typically have millions of splats. That's a lot of data, so it's no wonder it can be pretty photorealistic. (Still impressive, though.)
The actual changes this led to are:
1) As you mentioned, they added noise. But notably, they state it was "designed carefully" to conform to the requirements of SGLD and they detail how they designed the noise.
2) They simplified the original operations of "move, split, clone, prune, and add" and their related heuristics, into a single operation type. They do so guided by existing knowledge about MCMC frameworks, leading to a simpler model with stronger theoretical underpinnings (huge win!).
3) Adjustments to how gaussians are added and pruned to better fit with the new model. This seems more like housekeeping rather than something novel in and of itself.
What is the practical difference here? MCMC itself samples more from higher probilities than lower ones (ie. towards a local minimum). Is it just that we sample more from lower ends of the distribution? Or is it more about formalizing the previous algorithm so that it is easier to play with the different parameters? (eg. the acceptance threshold)
Speaking of uses for gaussian splatting - would anyone have any unique use cases for being able to generate splat models wildly faster?
When you get used to clicking a few links to to the a cited reference the whole scroll-to-reference, copy-paste to google, scroll to paper song and dance seems a lot less fun.
Probably a cool paper though.
Edit: The TeX source has the `draft` option for hyperref that disabled hyperlinks in the produced pdf. External links to references aren't recoverable from the included main.bbl (probably because it was built with the `draft` option).