WildGaussians: 3D Gaussian Splatting in the Wild
arxiv.org
arxiv.org
Truly an outlier ;)
This implies pretty obvious & severe limitations that it's not modelling where lights are or how things bounce, but it is really neat to see & works shockingly well for this use case of moving forward through a scene.
All these in-the-wild-methods share a similar setup, in that they do two things: appearance modelling (for daytime and weather and season and exposure, usually with per-image and per-Gaussian embeddings), and modelling/masking of transient objects (for tourists and dogs and water and blimps).
WildGaussians (this post): For the appearance modelling, they have a learnable per-Gaussian and a per-image embedding. They are fed into an MLP to produce affine transformations applied to the spherical harmonics (SH) params. So if you like a fixed appearance, you can pre-compute SH params and throw the scene into a standard 3DGS renderer.
The affine transformations are inspired by "Urban radiance fields", which predicts affine params form image embedding alone. WildGaussians use also per-Gaussian appearance embedding for local changes.
For the transients, they have an "uncertainty modelling" module, which computes DINOv2 features of the rendered image, and of the GT image. They compare, upsample and binarize them into a per-pixel mask, which is thrown onto the DSSIM and L1 loss.
Paper reads well, probably interesting to dive into how the uncertainty modelling really works. Straightforward setup (with the split into appearance and uncertainty), can be followed along well.
SWAG: appearance modelling somewhat similar to WildGaussians. Has image embedding, looks up per-Gaussian embedding from the Gaussian coords in a hash grid (hello Instant NGP). Feeds them together with the Gaussian color into an MLP. Which produces image-independent color and an image-dependent "opacity variation". So instead of masking out transient stuff in the image loss during training, they learn for each Gaussian (through hash grid), whether it's visible in a particular image.
The authors note that they also considered using "Urban Radiance Fields" affine transformations, but that affine colors cannot model all appearance changes.. which is why WildGaussians have the per-Gaussian embedding, I think?
Interesting that they can reproduce per-image transient objects. But maybe the static stuff looks a bit worse from that (look at the water in the Trevi fountain, its.. missing).
Bit hard to quickly follow along with the opacity variation stuff, this would take a bit of time to grok. But overall also quite straightforward setup, interesting read.
Wild-GS: Global appearance embedding (per image), per-Gaussians local reflectance, material attributes per Gaussian. Fusion network decodes SH from these three components. For some reason projects the points from the image from the depth into 3D, and looks that up in a triplane, instead of looing up the triplane from the Gaussian position? There's "2D Parsing" and "3D Wrapping", and it's quite convoluted, I'd need more time to understand what's going on here.
Gaussian in the Wild: extracts image feature with a UNet, reshapes into a bunch of feature maps (K feature maps + projection feature map?). The Gaussians sample from these feature maps. Features are fused with an MLP into a color. Some adaptive sampling is apparently required.
Transient objects are handled by a 2D visibility map obtained from a UNet as well.
I think the main idea is to train networks that can extract information from the input image to model the appearance and transients. This is different from WildGaussians and SWAG, which train per-image and per-Gaussian (directly or thourgh 3D lookup) embeddings, and only small decoder MLPs.
WE-GS An In-the-wild Efficient 3D Gaussian Representation for Unconstrained Photo Collections: may be similar to Wild-GS. Too much stuff going on, to understand from quick browsing. If you like the <2D input image feeds into lots of different networks> idea (like Wild-GS), you may want to read this one, too.
Robust Gaussian Splatting: tackles motion blur by modelling the camera poses as a Gaussian distribution. And defocus blur (from physical apertures) through an additional covariance on the Gaussians. Also have an RGB decoder function with per-image embedding for some appearance modelling (different exposures of the same scene). Interesting to read when you want to get rid of motion and defocus blur. For an in-the-wild appearance modelling, choose one of the other methods, it's very simple here.
SpotlessSplats: Ignoring Distractors in 3D Gaussian Splatting: spatio-temporal semantic clustering of objects that are likely transient (moving dog, "transient distractor in casual capture"). Also offers a new densification/pruning scheme. Focus is on casual capture, no appearance / in the wild stuff. Results look great, probably an interesting read!
The idea is that by layering enough of these on top of each other, you can make a high quality visual representation not bounded by geometry.
Think of them like brush strokes in mid air versus a sculpture.
There has been a big explosion of interest in the area since this release of https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/ which proposed a way to generate the point cloud data from photographs in a way that is procedurally a lot like photogrammetry. Hundreds of research projects have spawned off of that paper.
That said, Gaussian splats are based on very old papers that do have use in Hollywood. The closest example is radiance caches for Renderman.
This would store points in space with very similar information to a splat, then interpolate between them to provide things like bounce lighting.
The first Pixar film to use this throughout was Up.
What are some examples of those representations?
They just have a higher up front cost to make than Gaussian splats but are infinitely more directable.
Ultimately the ability to control everything at a micro and macro level is what film wants, not cheap rendering.
Gaussian splatting represents a scene a cloud of 3d gaussian ellipsoids, with direction-dependent color components (usually represented using spherical harmonics) to deal with effects like reflections. The "Gaussian" part is important, because gaussian distributions are easy to differentiate, making it possible (and fast) to optimize the positions, sizes, orientations, and colors of a collection of Gaussian splats to minimize the difference between the input photos and the rendered scene. This optimization is usually done by starting with the same 3d point clouds and camera poses estimated using the same or similar tools as traditional photogrammetry (e.g. COLMAP), and using this point cloud to place and color your initial Gaussian splats. One of the key insights in the original Gaussian splatting paper was the use of some heuristics to determine when to split a splat into smaller ones to provide higher detail over a given area, and when to combine splats into larger ones to cover uniform/low detail areas.
The nature of Gaussian splats being essentially fancy point clouds means that they can't currently be easily integrated into existing 3d scene manipulation pipelines, although this is rapidly changing as they gain popularity, and tools to convert them into textured meshes and estimate material properties like albedo, reflectance, and so on do exist.