CityGaussian: Real-time high-quality large-scale scene rendering with Gaussians
dekuliutesla.github.io
dekuliutesla.github.io
EDIT: here it is, and I was right! https://city-super.github.io/matrixcity/
> Binding New 3D Gaussians to the Mesh
> This binding strategy also makes possible the use of traditional mesh-editing tools for editing a Gaussian Splatting representation of a scene
No images involved, so no training required.
Basically the problem with sparse pictures and point clouds in general is their lack of topology and not precise spatial position. But when you already have the topology (eg with a mesh), you can extract (optimally) a set of points and compute the radius of the splats such that there are no holes in the final image (and their color). That is usually done with the curvature and the normal.
The 'optimally' part is difficult, an easier and faster approach is just to do a greedy pass to select good enough splats.
Through pixel perfect alignment? Resolution?
The matrixcity map is different, but is somewhat similar to that of the map in Matrix Awakens. You can see from the design breakdown on [3] this page by a technical lead on the Matrix Awakens project.
edit. If you look further into the [4]github codebase, under the MatrixPlugin section it explicitly states that they used the city-sample project.
[1] https://www.unrealengine.com/marketplace/en-US/product/city-... [2] https://www.unrealengine.com/marketplace/en-US/learn/city-sa... [3] https://quentinmarmier.artstation.com/projects/xYeKNO [4] https://github.com/city-super/MatrixCity
Example 1 with code linked: https://twitter.com/kfarr/status/1773934700878561396
Example 2 https://twitter.com/3dstreetapp/status/1775203540442697782
getting errors on that first link in devtools
Uncaught (in promise) Error: Failed to fetch resource https://tile.googleapis.com/v1/3dti...
Code is available via glitch url
Would love to follow. But not, you know, over there.
Real-Time if you have $8k I guess.
“What a time to be alive” indeed!
I was doing "realtime ray tracing" on Pentium class computers in the 1990s. I took my toy ray tracer and made an OLE control and put it inside a small Visual Basic app which handled keypress-navigation. It could run in a tiny little window (size of a large icon) at reasonable frame rates. Might even say it was using Visual Basic! So yeah "realtime" needs some qualifiers ;-)
Moore's law may be dead in general, but computing power still increases (notwithstanding the software bloat that makes it seem otherwise), and it's still something to count on wrt. bleeding edge research demos.
The issue is that in Computer Science "real-time" doesn't just mean "pretty fast", it's a very specific definition of performance[0]. Doing "real-time" computing is generally considered hard even for problems that are themselves not too challenging, and involves potentially severe consequences for missing a computational deadline.
Which leads to both confusion and a bit of frustration when sub-fields of CS throw around the term as if it just means "we don't have to wait a long time for it to render" or "you can watch it happen".
I think that pretty much meets the definition of "you can watch it happen".
Essentially there is real-time systems and real-time simulation. So it seems that they are using the term correctly in the context of simulation.
Consumer GPUs are probably 2-3 generations out from being as capable as an A100.
You can convert it to a mesh, but in the process you'd lose the quality and realism that makes it interesting.
By the way, what's the compute power difference between an A100 and a 4090?
So for this specific application it really depends on where the bottleneck is
Check https://github.com/pierotofy/OpenSplat for something you can run on your 10 year old laptop, even without a GPU! (I'm the author)
They worked with a basic 970 GTX on a big 3d screen and also on oculus dk2.
https://www.techpowerup.com/forums/attachments/all-cards-png...
So I don't know see an insurmountable problem.
I think generating traditional geometry and materials from gaussian point clouds is maybe interesting. But photogrammetry has already been a thing for quite awhile. Trying to render a giant city in real time via splats doesn't feel like "the right thing".
It's definitely cool and fun and exciting. I'm just not sure that it will ever be useful in practice? Maybe! I'm definitely not an expert so my question is genuine.
Traditional photogrammetry really struggles with complicated scenes, and reflective or transparent surfaces.
Current photogrammetry to my knowledge requires much more data than NeRfs/Gaussian splatting. So this could be a way to get more data for the "dumb" photogrammetry algorithms to work with.
100+FPS in browser? https://current-exhibition.com/laboratorio31/
900FPS? https://m-niemeyer.github.io/radsplat/
We have 3 decades worth of R&D in traditional engines, it'll take a while for this to catch up in terms of tooling and optimization but when you look where the papers come from (many from Apple and Meta), you see that this is the technology destined to power the MetaVerse/Spatial Compute era both companies are pushing towards.
The ability to move content at incredibly low production costs (iphone movie) into 3d environments is going to murder a lot of R&D made in traditional methods.
Edit: I was in low power mode, it runs quite smoothly
https://medium.com/@heyulei/capture-images-for-gaussian-spla...
For 3D gaming engines? I struggle to see how the fundamental primitive can be made to sing and dance in the way that they demand. People will try, though. But from this perspective, gaussians strike me more as a final render format than a useful intermediate representation. If they are going to use gaussians there's going to have to be something else invented to make them practical to use for engines in the meantime, and there's still an awful lot of questions there.
For other uses? Who knows.
But the world is not all 3D gaming and visual special effects.
Many of the core papers for this came from Meta's VR team (codec avatars), Apple ML (Spatial Compute) and Nvidia - companies deeply invested in VR/Spatial compute. It's clear that they see it as a key technology to further their interests in the space, and they are getting plenty of free help:
After being open sourced in May last year, there were 79 papers overall published on the topic.
It's more than 150 this year, more than one a day, advancing this "dead end" forward.
A small selection:
https://animatable-gaussians.github.io/ https://nvlabs.github.io/GAvatar/ https://research.nvidia.com/labs/toronto-ai/AlignYourGaussia... https://github.com/lkeab/gaussian-grouping
In the meantime, if it isn't, it will hardly be the first promising new graphics technology to turn out to be completely unsuited for all the things people hoped for.
Most of what you linked to appears to correspond to what I intuitively described as them being an output format rather than useful directly; the last paper appear to go in the other direction to extract information from them but again doesn't function on the splats directly. The actual work isn't being done in the gaussians themselves, and the interesting results are precisely in what is not being done through the splats... but pointing that out explicitly that's not how you get funding nowadays. Two otherwise-identical proposals, but one that sings the praises of the buzzwords while the other is phrased to be critical of it, will have very different outcomes.
After testing it, if fails in very basic cases. And it is normal that it fails, non Lambertian materials are not reconstructed correctly with SfM methods.
I might be misunderstanding what you're trying to say. Could you elaborate?
However, non-Lambertian points move non linearly in viewing space (eg a specular point depends on the viewer pose).
So, automatically, their positions in space will be false, and you'll have floating points.
Gaussian 'splats' may have the potential to render non-Lambertian stuff using for example the spherical harmonics (even if I don't think the viewer use them if I'm not mistaken). But, capturing non-Lambertian points is very difficult and an open research problem.
Is it? So far it seems like the storage size is massive and the detail is unacceptably low up close.
Is there a demo that will make me go “holy crap I can’t believe how well this scene compressed”?
The key is not to compress but to leverage the property of neural radiance fields and optimize for entropy. I suspect NERF can yield more compact storage since it's volumetric.
Not sure what you mean by "unacceptably low up close". Most GS demos don't have LoD lol.
When the camera gets close the "texture" resolution is extremely low. Like, roughly 1/4 what I would expect. Maybe even 1/8. Aka it's very blurry.
In practice the answer to will this be useful is yes! Subdivision surfaces coexist with nurbs for different applications.
Heck, you can even train one from scratch in a minute on an iPhone [1].
This technique has been around for less than a year. It's only going to get better.
And? It's always going to be even faster to not have lighting at all.
https://en.wikipedia.org/wiki/Monte_Carlo_(disambiguation)
8 entries in "Science and Technology" alone.
I was more thinking you'd run this tool, and then have an algorithm convert it( bake the mesh).
For hybrid pipelines to work the splatting algorithm would probably need to output the standard G-Buffer channels (unlit surface color, normal, roughness, etc) which can then go through the same lighting pass as the triangle-based assets, rather than the splatting algorithm trying to infer lighting by itself and inevitably getting a result that's inconsistent with how the triangle-based assets are lit.
Think of those old cartoons where you could always tell when part of the scenery was going to move because the animation cel would stick out like a sore thumb against the painted background, that's the kind of illusion break you would get if the lighting isn't consistent.
It is not too difficult to go to a 2D normal field over the 3D gaussians..
The objects represented in the point cloud have no inherent metadata embedded (ie its a chair, table, person etc) so any kind of interaction is super hard.
Its not impossible, but currently not practical.
More over its not that optimised for real-time rendering. Yes, a lot of points have been pruned, but its far more optimal to have lower resolution meshes
Am I missing something?
And you might want it to be the motion relevant to the camera. For Mario, probably not, but for an FPS you want to edges of the screen to blur as the camera moves forward.
* Virtualised geometry (Nanite) allowing very detailed models
* Very high quality models and textures from photogrammetry (Megascans)
* Real-time global illumination (Lumen)
Combining these is what allows the very high fidelity demos, as they’re each step changes from the previous techniques in Unreal. Megascans (and the Quixel library) are a big part of the “photorealness” of these demos, because they’re basically literally photos.
I can't wait to see once it's applied to the world of driverless vehicles and AI!