Infinite Photorealistic Worlds Using Procedural Generation
arxiv.org
arxiv.org
Here's a memorable intro, from 2009: https://news.ycombinator.com/item?id=31636482
Earlier than that, another famous 64k from 2000:
See for example some 64kB demos that have landscapes / nature:
- Paradise by Rgba, 2004 https://www.pouet.net/prod.php?which=12821
- Gaia Machina by Approximate, 2012 (https://www.pouet.net/prod.php?which=59107)
- Turtles all the way down by Brain Control, 2013 (https://www.pouet.net/prod.php?which=61204)
Of course, it's not exactly the same, the demoscene tends to focus on a few carefully crafted scenes and do everything in real-time (with up to 30s of precalc), but I think there's a lot of similarity.
The demo scene was already standing on the shoulders of giants.
https://www.youtube.com/watch?v=BFld4EBO2RE
https://www.youtube.com/watch?v=8--5LwHRhjk
EDIT: To be clear, I was not being flippant, I was assuming that there is some nuance to how "assets" is used here that I'm not aware of since I'm not a 3D artist which would make this work novel compared to the work in the demoscene. Beyond being a general purpose asset generator, which is indeed very impressive but not what is being argued here.
What is being argued here? I don’t see where this turned into a competition.
It’s a paper presenting a system of new algorithms that can generate geometry, textures and lighting with extreme photorealism. It seems to do that very well. It’s not a study of the procedural 3D graphics evolution, pretending to be an entirely new technique or the first to do it.
So if they don't view the demoscene as relevant previous work to their paper then I want to understand why they think their work falls under a different kind of procgen.
Again, not a contest. The paper is exploring new techniques for generating content. Why all the negativity?
No new ground is broken here in rendering techniques, nor in procedural generation, nor in actual content. AI is generating actually photorealistic content today, AI-adjacent techniques such as NERFs as well as traditional but sophisticated rendering like UE5 are the leading edge of rendering. If you're going to make a strong claim, you should deliver something.
I honestly don’t know what you’re reacting to, what strong claims? “Infinite” is a given for procedural assets, and is highlighted in comparison to the existing partially asset-based generators (they go over existing software on page three), and the output certainly qualifies as photorealistic. There are no claims in the paper about being the first to do anything.
The sheer scale of what this does, how general it seems to be (instead of a single special-purpose animation like in the demoscene), as well as the fact that the output is structured and labeled assets I would consider novel, and very impressive.
"Ours is entirely procedural" == ""Infinigen is entirely procedural"
"relying on no external assets" == "every asset, from shape to texture, is generated from scratch via randomized mathematical rules, using no external source"
If that doesn't make it make it clear, could you elaborate on the part that doesn't click and I'll try and explain further.
The quality is good and it's giving you all of the maps (as well as the .blends!). It seems great for its stated goal of generating ground truth for training.
However, it's very slow/CPU bound (go get lunch) so probably doesn't make sense for applications with users behind the computer in the current state.
Additionally, the .blend files are so unoptimized that you can't even edit them on a laptop with texturing on. The larger generations will OOM a single run on a reasonably beefy server. To be fair, these warnings are in the documentation.
With some optimization (of the output) you could probably do some cool things with the resulting assets, but I would agree with the authors the best use case is where you need a full image set (diffuse, depth, segmentation) for training, where you can run this for a week on a cluster.
To hype this up as No Man's Sky is a stretch (NMS is a marvel in its own right, but has a completely different set of tradeoffs).
EDIT: Although there are configuration files you can use to create your own "biomes", there is no easy way to control this with an LLM. Maybe you might be able to hack GPT-4 functions to get the right format for it to be accepted, but I wouldn't expect great results from that technique.
Out of curiosity, what kind of cpu / ram are you meaning here?
Asking because I have some spare hardware just sitting around, so am thinking... :)
Lots of places have servers with 128-256GB of ram around though.
I have a decent PC (AMD 3990X 64-Core Processor with 256 GB of RAM), I'd have installed better/more components but that seemed to be the best you could do on the consumer market a few years ago when I was building it.
Are they using the same RAM I'm using with a different motherboard that just supports more of it? Or are they using different components entirely?
Apologies for what I'm sure is a very basic question but it would be interesting to learn about.
Here's what a lowly 256GB server looks like. For a TB just imagine even more sticks:
The only real exception to that would be for database or reporting servers, which sometimes might have higher ram requirement (eg 384GB).
That's pretty much it though.
I gave a standing ovation when that popped up in the video.
What a gift. Gotta love it.
I think one could make a case for the former, but the latter is subjective (I liked the game, although it is a bit shallow).
Then add spaceships so you can actually travel between them.
Godot runs in the browser these days, so that could work.
"Each part generator is either a transpiled node-graph, or a non-uniform rational basis spline (NURBS). NURBS parameter-space is high-dimensional, so we randomize NURBS parameters under a factorization inspired by lofting, composed of deviations from a center curve. To tune the random distribution, we modelled 30 example heads and bodies, and ensured that our distribution supports them."
This strikes me as a fairly random approach. No wonder why those creatures look the way they do. I fail to see why this is worth a scientific paper as it appears to be no more than a student project with a number of contributors across different fields. Building a (somewhat) procedually based asset library has been done countless times before by game dev studios big and small.
" The wall time to produce a pair of 1080p images is 3.5 hours."
Ouch. It's in Python, but still...
This only works if the data contains such redundancy. That depends on how it's generated. Not clear if this type of generator creates such redundancy. Something using random processes for content generation won't do that.
Nanite is optimized for large areas of content generated by a team of people working together, as in a major game project. There's a lot of asset reuse. Nanite exploits that. It's not a good fit for randomly generated assets, or for assets generated by a large community, as in a serious metaverse. The redundancy isn't there to be exploited.
Many of these artist-made high-fidelity meshes make heavy use of Z-brush style sculpting and procedural materials, which have similar characteristics to these randomly generated assets.
I'd kind of expect that most of the core stuff is being done in C++ or similar behind the scenes though. Maybe (haven't looked). :)
Eventually we concluded that machine learning wasn't a good fit for our problem, and our users were very keen to maintain that conclusion.
This makes the repo much more useful to me.
A nice next step or addition would be to take the results and remesh them to lower poly models so it can be used in a game engine to walk around in.
That was around for decades.
https://m.youtube.com/watch?v=Ws1zpk9GWkk
Much, MUCH more limited, but pretty cool.
Of course AI input would work well with this...