Gaussian splatting is pretty cool
aras-p.info
aras-p.info
then..
> And finally, they have resisted the temptation to do “neural” anything ;)
So, they are doing something similar to NeRF, in a sense, but using a different basis function, and a slightly different target. The whole "neural" or not "neural" is just about what you are optimizing, but it's not that different, conceptually -- they are optimizing a data-driven approximator of a 3d scene-related variable based on images.
The big difference of course, based on reading this, is that NeRF models the whole light transmission function ("radiance field") whereas this seems to model only the boundary conditions (what the light "hits"). So given the different modelling target, then yes, a different, and perhaps simpler, representation basis definitely seems warranted, and is shown here to give good results. But the comment, "they avoided neural anything" feels a bit smug, as if "avoiding the hype" is a laudable goal on its own merits. As if people are using neural networks for no reason but because they're cool, and not because they're actually an appropriate solution that are shown to work really well in practice. They surely are cool, don't get me wrong, but the hype is so often justified, because neural networks are really good function approximators. They are also not just one thing -- the choice and breadth of architectures that we happen to call "neural networks" is huge, so explicitly avoiding this huge potential solution space for some reason without good reason feels pretty specious; heck you could even consider having a Gaussian "output layer" to a neural network, which is called an RBF, being something like this "splat" approach. (This is in response to the blog post, not the paper -- I'm sure the paper has plenty of justification for modeling choices.) Having said that, of course it's interesting to also explore simpler, perhaps more efficient approaches, but I don't see "resisting using neural anything" as a good thing, just because they are popular. It ignores that they are popular for a reason.
(nb: This assessment is just based on the blog post and giving my impression from the tone of the introduction, I haven't read the paper yet, so might not be accurate with respect to the paper contents. I just found this off-hand comment in the post worth reflecting on.)
It’s just a function call and you can always transform the NeRF into a new representation if you want to use traditional 3d software.
Transforming it to a mesh or SDF is cheating, then it's no longer a NeRF and of course the same cons doesn't apply anymore.
AI = means you don't understand how to quantify the problem
AI is a great hack to quickly and unknowingly solve classification problems and other similar things. But you trade massive datasets for explainability. And that's a pretty bitter pill to swallow.
Even the "simple" AI's with only a few TBs of data for training still have low/no understanding what's going on to identify the features and how they relate. Sure, we can tell this neuron identifies this squiggle, but what does it mean? No fucking clue. It also shows us collectively that most of these "hard" problems we have no clue how to understand them. And although AI allows us to "do it", it's a bad hack for actually understanding, optimizing, and doing.
> That it works without you understanding it is a failure of your understanding not of the approach.
Cool. Then point me towards the tools to explain how weights in a moderate-ly sized NN are created with training data, and associated with each other and how they find the features in question.
Oh yeah. Those tools don't exist.
It's not at all clear that there is or even should be a deeper explanation than "a function with sufficiently many parameters can approximate any other function with only a small change to those parameters, and thus allows tractable solving" - whether by gradient descent or some hokey approximation our neurons are probably doing.
We have even less of an idea how our own system of "understanding" works than we do of neural networks "understanding". If you're going to be mad at tools that don't provide satisfying explanations for how they work, then look no further than your own brain. Are you going to give up on thinking, just because when you look at it closely you realize you have no idea how on earth you're even doing it?
One of my personal favourites is Zeiler and Fergus https://arxiv.org/abs/1311.2901
Tom Zahavy’s “Greying the black box” on deep RL comes to mind too https://arxiv.org/abs/1602.02658
OpenAI has their Microscope tool for investigating individual neurons https://openai.com/research/microscope; more recently OpenAI shared research and tooling leveraging a powerful language model to explain a smaller model’s neurons. https://openai.com/research/language-models-can-explain-neur...
For a true toy box exploration playground.tensorflow.org is worth a play
These may or may not satisfy your curiosities — your perfect tool may not exist…yet. I would argue that the sampling of extant work above are indicative that it and more are possible.
If you look at one of the most widely used text books machine learning is one of seven chapters, and not all of that is deep or supervised learning.
In my view calling a solution "neural" and "AI" because you are optimizing a stack of relus, and then turning around and solving the same (or almost the same) problem by optimizing a bunch of Gaussian model parameters on the same data, and then saying that is "not AI" and "at least it's not neural!" is just a very specious and unnecessary way to delineate things imho. They're just two solutions to the same problem, both data driven, both generating "models".. they have differences, just not differences that I think warranted such a comment in the blog, and moreover I thought it weird to attribute some kind of normative evaluation based on how you categorize the solution.
In this instance, I’m going to choose to take an optimistic stance and believe the authors were talking a friendly elbow jab rather than fueling a mini culture war. The researchers leveraging DNNs are doing great work, but perhaps could be wrangled in just a bit :)
The title of the paper is "3D Gaussian Splatting for Real-Time Radiance Field Rendering". Each rendered pixel weights the contribution of unbounded view-dependent Gaussians. So, no, that's not a difference.
I read Aras' quip as more narrowly technical. OFC it's ambiguous and you'd have to ask the man himself.
The gist of NeRF is to obtain an NN representation of a 5D light field (3D position + 2D direction) from samples (photographs) of the real-world light field. Alarm bells ring already- 5 dimensions isn't that many! Considering NeRF has always used a low-rank spherical harmonic representation of the directional domain, it's even more like 3D-and-change. To reconstruct a function of such low dimensionality, why choose an NN?
Then at inference time, for each pixel, you have sample the NN repeatedly over the view ray. This part is exceedingly silly, as compact representations of light fields are a solved bread-and-butter problem in graphics.
Later on Plenoxels explicitly took the "Ne" out of "NeRF", giving far higher training and inference performance (also mentioned ITT). To be fair, and later still, Nvidia somewhat redeemed NNs here with Instant NeRF: https://nvlabs.github.io/instant-ngp/assets/mueller2022insta... ...where the twist was to interpolate fancy input emeddings, which are run through a tiny NN. That tininess is important, as the need to fetch NN weights from VRAM would kick NNs right off the Pareto frontier.
Zooming out, NNs have only seen wide adoption in graphics engineering for reconstruction from sparse data (inc. denoising). Makes sense, as that's a high-dimensional problem. Still, beware that the NN solutions rarely blow handmade algorithms out of the water. I also think using tiny NNs for compression- closely related to reconstruction- has a future too. Beyond that, if NNs were to set the graphics world ablaze, it would've happened by now.
Lots of graphics engineering is just approximating functions, so it's natural NNs have some place here. However, our functions tend to be more understandable, tractable, malleable. It's not an application domain where it's virtually impossible to write an algorithmic solution by hand (let alone one that performs well), like natural language understanding.
As a bit of tangent (but wondering whether someone can answer) - the article also makes mention of point based rendering and indeed the fact is has been a staple of particle systems for a long time. However, especially with recent games, I have noticed (purely subjectively) a very subtle shift to a new style of particle systems which are on the one hand fully point oriented (compared to (textured) fragments) but on the other behave more like a physics systems.
Examples:
- Hogwarts (heavily): https://www.gamespot.com/a/uploads/original/1816/18167535/40... - Forspoken (heavily): https://oyster.ignimgs.com/mediawiki/apis.ign.com/project-at... - Starfield (though more rarely): https://dotesports.com/wp-content/uploads/2023/08/temple-loc... - AC6 - FF16 (heavily)
It's more obvious when you see it 'in motion'. The common denominator seems to be particles as colored transparent points with physics. Especially on console systems it seems that developers are using this for very cheap (CPU-wise, all on GPU) effects.
Anyone in gamedev who has some insight in this?
I don't think they're "from" that, but that's the first landmark I can point to.
Which in some ways I guess makes sense, as the shift towards AI has heavily pushed GPUs almost towards a "sizeable-reduced-instruction-set-GPU-in-GPU" approach with RT cores/tensor cores, etc.... ;P
Side note as well, in/from my experience at least, I think these systems may not be as hard to render as it is for the author. A few years ago, someone (at NVIDIA I think?) wrote kernels to move the SE(3) kernels to tensor cores, so I wouldn't be shocked if some of that could be ported to the spherical harmonics portion of the Gaussian splatting during both compression and runtime: https://developer.nvidia.com/blog/accelerating-se3-transform...
Also, side side note, gaussian splatting should be quite efficient...I think? Due to technically always having support in 3D space (and hopefully not too much of a problem with having good support in 3D space). This should mean that even 'sloppy', quick-conversion calculations should work pretty decently in the end.
I say all of this knowing very little about most optimizations like billboarding, how things like nanite work, etc, etc. I do like it tho! ;PPPP
Without a fluid sim, point-based particle systems just look like fireworks. It was a cool effect in, like, the early 90s, but is is passe today.
The next step up from that is having each particle move independently with its own physics (some momentum, maybe a little wandering around from "wind", etc.) but then rendering them using little texture billboards. That's what games did up until relatively recently and looks pretty good for explosions, smoke, etc.
But now machines are powerful enough and physics algorithms clever enough to actually do fluid simulation in real-time where the particles all interact with each other. I think that's what you're seeing now.
Edit: I also just found this: https://www.youtube.com/live/569oSOSoKDc?si=8V5buRMoI3IKqLQp... -- which is very close to what you describe and fully matches the kind of particle systems I was hinting at, thanks!
I'm assuming Unreal Engine has something similar, but I don't much experience with it.
Way back during the transition from Doom to Quake, which in some ways really marked the transition from 2D to 3D, Quake's particle systems also relied overwhelmingly on small flat colored particles and their motions, rather than larger textured sprites and their silhouettes. (Quake did use a few sprites, but it was few and far between).
And I think the reasoning was pretty straight forward even back then; in a 3d game world, there are a lot of conceptual and architectural benefits to only working with truly 3d primitives - and point sprites often can be treated like nearly infinitely small 3d objects.
Whereas putting 2d sprites into a 3d scene introduces a bunch of kludges. In particular, 2d sprites with any partial transparency need to be sorted back to front in a scene for certain major blend modes, which gets really troublesome as there are more and more of them. They don't play nice with zbuffers. And because they need to be sorted, they don't always play nice with the more natural order you might prefer to batch drawing to keep the GPU happy. And likewise, they have a habit of clipping into 3d surfaces in ways that reveal their 2d-ness. There's probably more things I'm forgetting.
These are all issues that have had lots of technical workarounds and compromises thrown at them over time, because 2d transparent textures have been so important for things like effects and grass and trees. Screen door transparency. Shaders to change how sprites are written into zbuffers. Alpha testing. Alpha to coverage. Various follow-on techniques to sand down the rough edges of these things and deal with aliasing in shaders. And so on.
And then there's the issue of VR (or so a cursory skim suggests). I haven't spent time doing VR development myself, but a quick refresher skim of articles and forum posts suggests that 2d image based rendering in 3d scenes tends to stick out a lot more, in a bad way, in VR than on a single screen. The fact that they're flat billboards is much more noticeable... which is roughly what I had guessed before I started writing this comment up.
All of those reasons taken together suggest why people would be happy to move on from effects based on 2d texture sprites in 3d scenes, to say nothing of the other benefits that come from using masses of point sprites specifically themselves (especially in terms of physics simulations and such).
Once the tools around this mature a bit more, I'm super excited to revisit those old rooms and be hit with a wave of nostalgia.
Radiance fields have no concept of light emission, reflection, absorption, etc. instead everything is mushed into one value: The light transported. In that sense radiance fields are just 3D photos.
You would have to perform reverse-rendering / photo grammetry and estimate where the light sources and the surfaces are, what materials they have and so on. Then you could use traditional path tracing methods on that again.
Another thing to think about might be videos (not animation): Continuously capture the radiance field over time and then try to compress away the similarities in between frames to gain temporal coherence.
Is depth not properly preserved or something?
Remember that it's a bit tricky to talk about depth when the gaussians have both a position (mean value) and a size (covariance). The bicycle spokes are made up of long thin splats, what value do you assign to one of those? That's why I think you would have to sample new points from them as a first step.
Here's an edge case I can imagine for dynamic lighting - Say you capture a scene indoors, and a table casts a dark shadow on the floor. But the NeRFs don't try to understand light sources and shadows yet, so it wouldn't know whether the floor is painted black, or a white surface shadowed by the table, or if there's actually a blue Stanford bunny hiding in the shadows.
The 3D scanning rigs that capture small objects like people's faces handle this by manipulating lighting and sampling the BDRF directly. If you can't manipulate the lighting, you can probably guess a BDRF, but there will be limits.
Re-animating might be easy. But capturing an animation, I think you'd need multiple cameras or you'd have to settle for guesswork, like a neural network that can hallucinate the hidden side of a person based on the fact that they're a person. If you point the camera at someone who is walking, you'll get a good view of them from one side, but when you wind the video back, the network won't know what their far side looks like at all.
A few years ago Intel had a project to capture an animated scene with multiple cameras. The pitch was something like, "Just film everything, and you can position the camera in post-processing." I think they wanted it for football games, but I never heard of it shipping. And again, multiple cameras. Matrix-style.
The data requirements might become massive, but there are ways to do the interpolation where it isn’t so bad. If a static scene is 2GB, you should be able to get to a rough time of day approximation in less than 16GB, which is renderable on modern GPUs.
Then it’s “just” a matter of spending several years optimizing it while waiting for H100s to become consumer grade devices.
The only hope is to measure actual photons hitting actual sensors, which is why Gaussian splatting looks so real to begin with.
is a better alternative from your perspective?
When moving the camera to a perspective distinct from the original one, the vision disintegrates into kind of three-dimensional pixel blobs, somewhat similar to the notion of blobs here - just a lot less polished, surprisingly, compared to this paper.
[1]: https://steelseries.com/blog/how-to-braindance-cyberpunk-207...
For me, a rendering guy, that’s great! The data used at render time is very simple and flexible. Simpler than triangles even when you get into non-trivial operations.
Classically, researchers were interested in just using the splats at all. So, they just assigned a single color to each one.
This paper assigns a spherical harmonic-based color sphere instead. That gives it the view angle -> color function.
There is a second paper focused on moving splats. It just uses solid colored splats.
You could instead associate albedo/specular/roughness from real time materials and do real time lighting. But, you’d have to figure out how to generate/capture those values.
P.S. Not sure where the insight came from, but I was living near Boston at the time, and had young children :)
[1]: https://xi.zulipchat.com/#narrow/stream/197075-gpu/topic/Gau...
(Although I am suprised that there's not more attention being paid to the extra spatial relationships that can be inferred from merely knowing something is a video. Surely it could be used as an extra constraint when inferring camera poses - you know the camera can only move in certain ways from one frame to the next?)
Can I ask why you say this? Camera motion models including Kalman filters of various kinds and constraints on derivatives etc, are absolutely used in SLAM and photogrammetry afaik, and have been for a long time.
This was a couple of years ago probably. Meshroom and Reality Capture. Maybe I was mistaken or maybe they've improved?
Can you point me to software that does use video for photogrammetry in the way you describe?
As far as I know, the basic approach to SfM hasn't changed much in the last decade or so. It boils down to image feature extraction (using something like SIFT feature vectors), then a heuristic matching process, then bundle adjustment and outlier rejection.
Is there something off the shelf that does this, or a library that can generate these point clouds so I can go about trying to fit simple geometry to them to reconstruct the scene?
Please forgive the noob questions; I've not worked with automated 3d stuff or photogrammetry before.
It's a pretty decent tool, although not sure if videos work well.
I once tried to build a mesh from an aerial video by extracting certain frames and the result was ok-ish
What does this mean? Why would scale matter in a 3D model when you can scale it however you want in two seconds?
Also they said 'to-scale', but none of this applies, since if you are doing some sort of photogrametry you are getting an accurate model of what you are photographing anyway, that's the entire point.
You can supplement images with lidar, stereo images, or IMU odometry to capture true scale information at the SfM stage to get (approximately) scale-correct models.
Alternatively you can learn what the real world tends to look like and estimate absolute depth from images alone, if you have enough data. This might be defeated by toy cows and doll houses, but work for some applications.
This is not a term that makes sense. It makes a model, there is no sense in having it scaled in any way except for being correct. Saying 'up to scale' is like saying 'feet running' or 'mouth eating' or 'face talking'.
the scale is arbitrary
Only if you don't measure anything. If you have the camera correct and know how high something is or just know some depth or distances photogrametry makes an accurate model. 'Scale model' is a term from when scaling the model down made it easier to make. In the computer, it just means it isn't accurate.
It's used when you don't know the scale.
You can estimate all the geometry 'up to' but not including a scale term.
> and know how high something is or just know some depth or distances
There you go: you just added in some extra information to constrain the scale estimate as I described.
I recently tested exactly this use case, with iPhone apps 3D Scanner, Polycam and magicplan.
The results were, how do I put it nicely, not very useful.
For example, the floor wasn't even flat/straight! One would think this would be a basic constraint (or “inductive bias”) of the algorithm…
I haven’t tested Luma yet; also my example wasn't literally the most basic one (square empty room), instead there was some clutter on the floor, irregularly shaped room, a chair in the middle of the room, and open doors.
Maybe you could say it's basically a 3D version of Alone in the Dark using ellipsoids instead of polygons as the primitive.
The quality of the animation is unusually good for the time, which is really what creates the coherent illusion. Wikipedia says the game had just one developer paired with a "film animation expert", which makes a lot of sense.
https://en.wikipedia.org/wiki/Ecstatica
I'm shocked I've never heard of this game until now.
There's coherence frame-to-frame of the artifacts used to render, but they may be 2D "watercolor" splotches, and some seem very manually seeded, e.g. the shapes used to fill in irises for eyes. In some scenes though the blobs seem to be 3D-represented.