Gaussian splats are much better suited for capturing things where you don’t have artists available and don’t have a ton of performance requirements with regards to frame time.
So things like capturing real estate , or historical venues etc.
It’s definitely not in the cards in the near term without a dramatic breakthrough. Splats have been a thing for decades so I’m not holding my breath.
What am I missing on the performance front?
It also ignores that everything else will be faster too then as well, and ignores needing to target different baselines of hardware.
Either way 5 years for a 3x improvement seems unrealistic. 4 years saw a little over a doubling of performance at the highest end with a significant increase in power requirements as well, where we’re now hitting realistic power limits.
Taking the 2080 vs 4080 as their respective tiers
153% performance increase 50% more power consumption 50% price increase.
So yes performance at the high end will increase, but it’s scaling pretty poorly with cost and power. And the lower end isn’t scaling as linearly.
On the lower/mid end (1060 Ti vs 2060 Super) we saw only a 53% increase in that same time period.
Here's splats from 2020 working at 50-60fps [1]. I think my overall point is I don't think it's performance that's holding it back in games but tooling & whether it saves meaningful costs elsewhere in the game development pipeline.
Otherwise flying cars will also be possible.
Also your splat is running in isolation. Any single system can run by itself at a good clip. That’s not indicative of anything when running as part of a larger system. Again, the discussion of performance is pointless without bounds.
This is not true I believe. There are plenty of papers out there revolving around dynamic/animated splat-based models, some using generative models for that aspect too.
There are also some tools out there that let you touch up/rig splat models. Still not near what you can do with meshes but I think fundamentally it’s not impossible.
With regards to dynamicism, there’s some papers yes but with heavy limitations. Rigging is doable but relighting is still hit and miss, while most complex rigs require a mesh underneath to drive a splats surface. There’s also the issue of making sure the splats are tight to the surface boundary, which is difficult without significant other input.
Other dynamics like animation operate at a very gross level, but you can’t for example do a voronoi fracture for destruction along a surface easily. And again, even at a large scale motion, you still have the issue of splat isolation and fitting to contend with.
The neural motion papers you mention are interesting, but have a significant overhead currently outside of small use cases.
Meshes are much more straightforward, and with advancements in neutral materials and micropolygons (nanite etc) it’s really difficult to make a splat scene that isn’t first represented as a mesh have the quality and performance needed. And if you’re creating splats from a captured real world scene, they need significant cleanup first.
Though I think matterport will just start using them since the other half of their product is the user experience on the web.
They’ll likely never be smaller than a mesh and texture though, because the data frequency will be higher. A wall can be two triangles and a texture. The same representation as splats will have to be many hundreds of points, roughly at the count of the pixels of the lowest resolvable version of that texture.
So I agree they’re far from optimal for data size. But they greatly reduce the complexity of data capture and representation.
It was not likely to be GS since there was tons of artifacts that didn't look like the ones GS produces, but they could have used it for such stuff.
For instance with some kind of 4D GS we could even remap the camera view entirely to have a virtual camera allowing us to see the shoot from the eyes of Steph Curry with Batum and Fournier double teaming him.
The current tech (Apple Vision Pro included) uses two photos: one per eye. If the photos were taken from a distance that matches the distance between your eyes, then the effect is convincing. Otherwise, it looks a bit off.
The other problem is that a big part of the 3D perception comes from parallax: how the image changes with head motions (even small motions).
Techniques that are not limited to two fixed images, but instead allow us to create new views for small motions, are great for much more impressive 3D photos.
With more input photos you get a “walkable photo”: a photo that you can take a few steps in, say if you are wearing a VR headset.
I’m sure 3D Gaussian splatting is good for other things too, given the excitement around them. Backgrounds in movies maybe?
Edit: others are mentioning real estate I'd think that will prefer some pre processing but ymmv
First if all, most GS take posed images as input, so you need to run a traditional photogrammetry pipeline (COLMAP) anyways.
The purpose of GS is that the result is far beyond anything that traditional photogrammetry (dense mesh reconstruction) can manage, especially when it comes to “weird” stuff (semi-transparent objects).
Not exactly. The "splats" are both spread out in space (big ellipsoids), partially transparent (what you end up seeing is the composite of all the splats you can see in a given direction) AND view dependent (they render differently depending on the direction you are looking.
Also - there's not a simple spatial relationship between splats and solid objects. The resulting surfaces are a kind of optical illusion based on all the splats you're seeing in a specific direction. (some methods have attempted to lock splats more closely to the surfaces they are meant to represent but I don't know what the tradeoffs are).
Generating a mesh from splats is possible but then you've thrown away everything that makes a splat special. You're back to shitty photogrammetry. All the clever stuff (which is a kind of radiance capture) is gone.
Splats are a lot faster to render than NeRFs - which is their appeal. But heavier than triangles due to having to sort them every frame (because transparent objects don't composite correctly without depth sorting)
https://en.m.wikipedia.org/wiki/Spherical_harmonics
Basically for each Gaussian there is a set of coefficients and those are used to calculate what color should be rendered depending on the viewing angle of the camera. And the SH coeffs are optimized through gradient descent just like the other parameters including position and shape.
I could also see a market for people who want to recreate virtual environments from old photos.
Also, load the model on a single-lens 360 camera and infer stereoscopic output.
The primary purpose of Gaussian splatting is to frontpage here every two weeks.
So I have been following GS tech for a while. I’ve not yet seen anything (open source / papers) that quite gets there yet. I do think it will.
In my opinion, there are two useful ways GS can bring to this industry.
The first is ability to use photo capture to re-render as a high production quality video similar to what people do with Luma AI today. While this is a really cool capability, it’s also not really that hard to do anymore with drones and gimbals. So, the experience of creating the same thing via GS has to be better and easier, and it’s not clear when that will likely happen due to how painful the capture side is. You really need good real time capture feedback to make sure you have good coverage. Finding out there’s a hole once you’re off location is a deal breaker.
The second is to create VR capable experiences. I think the first real useful thing for consumers will be so you can walk around in a small three or 4 foot area and get a stereo sense of what it’s like to be there. This is an amazing consumer experience. But the practicality of scaling this depends on VR hardware and adoption, and that hasn’t yet become commonplace enough to make consumer use “adjacent possible” for broad deployment.
I could see it being used on super high end to start out.