Nvidia tool generates full 3D models from a single still image
blogs.nvidia.com
blogs.nvidia.com
The horses did it for me. It made it look like some low effort parody.
So the comparison point should probably be like grey boxes or whatev people have been using as placeholders/throwaways so far, and compared to that these generated models actually look pretty great
I’m not sure why they chose to show off the “texture” view without even explaining what it is.
Also, while it may look iffy right now, it seems they’ve done the heavy lifting. A round of polish and in a couple years nobody will be laughing at this. (I’m already impressed, personally).
The heavy lifting is the not-yet-done fine details in the 3D-models. Vaguely car-shaped blobs not even respecting the original general shape of the car (e.g. the Cobra or the Toyota SUV) but with the correct texture extrapolated is nifty, but nothing more.
It sounds like they expect you to use it when you just need a vaguely correct shape anyway, at least as a starting point for further refinement.
They disabled comments on the article, and on the youtube video. Combined with YouTube's newly hidden dislike ratings, we can't know how many viewers are unimpressed. Hype has more breathing room than ever before.
[1] https://www.youtube.com/channel/UCbfYPyITQ-7l4upoX8nvctg
A pro could do better, but a) do you have a pro available? and b) can they afford the time to model all of the item you want to throw into your mockup?
https://en.wikipedia.org/wiki/Procedural_generation
.kkrieger got a lot of attention for using this method and only being 96KB.
Certainly, there are similar things in user space where you can share your creations from some games using a code that was generated or similar.
As to faces, Skyrim does that. There aren't unique face meshes for every NPC, there's some base meshes and various parameters to adjust texture and shape.
The latter sounds nice but in practice game execs just use that to cut corners and you end up with an unrewarding grirnd fest of a game.
Yes.
> I would rather have a 300 GB game
I, on the other hand, would not.
For example, increasing background details (more pedestrians, cars, traffic, sounds, etc) in a city scene can greatly add to the atmosphere regardless of your specific gameplay loop.
Going even further, things like GPT can generate text given an input. It's beyond me but couldn't something similar be used to describe environments on the fly to be generated, or generate NPC stories on the fly given inputs?
I think a lot of this tech is coming together and games are about to get a whole lot larger, more immersive, and cheaper to produce. Even just replacing voice actors would be a huge leap.
Present games just completely ignore file size because compared to GPU 1tb SSD are cheap and multi TB HD cheaper.
There are games that completely waste their space, but we can’t lump them all together. Some create vast detailed and beautiful worlds from their storage budget. Some of them, however, are large without any compelling reason to be, with many gigabytes of wasted space on assets that are barely used.
If you were to generate faces "on the fly" as a level streamed off disk then that would likely take up too much of your frame budget and your game would grind to a stuttering halt. The other way of doing it (and this is generally how games which use procgen work) would be to either have a precompute step when you first install or launch a game. This might give you a long install or launch time and would still mean that you have a huge install size, but you would avoid taking up a chunk of your frame budget. Or you could do some clever scheduling and level design where you force the player through "tunnels" between areas that gives you enough time to generate a few new faces and load in assets before they enter an area. This is a pretty common technique where you need to load in a bunch of assets or do some heavy calculatuins. The Witcher 2 had an interesting and slightly buggy version of this where whenever Gerald opened a door to go from outside to inside the camera would swing around so you couldn't see inside the building, a second long animation would play and then when the camera turned back so you could see into the now fully loaded interior if the building. It didn't work properly but it was a good idea.
You could do what ms flight does and have a locally installed low resolution world and a high resolution streamed world. Adapting this you could stream in procgen'd faces and other items generated offline. I think this would work rather nicely.
Lastly, I know it was just an example but still, faces are probably the last thing you want to procgen with out passing the results by a human filter as humans are incredibly sensitive to them looking "off". If all your characters end up looking grotesque and off putting you might be shooting your self in the foot. See the furore about eFootball's terrifying crowd models.
It also solves the problem of "we have to choose what objects to model because our 3D artists don't have infinite time". Now the artists review GAN output and tweak what the algorithm gets wrong, creating tons of content without needing to create every vertex from scratch. Very exciting
Also, creating new content from GAN might not be super appealing for the artistic qualities of the game. Why not do what Escape from Tarkov seem to do with photogrammetry on real objects? Many props like soup cans, couches, whole environments, weapons, can be acquired in real world reasonably easily.
..it worked, but the output *.3ds files were low-resolution and the textures only looked good from one direction - and the material's lighting was way off.
So the concept is nothing new, but it's never as simple as taking even a bunch of photos, let alone 1 photo - there's so much information necessary for future renders that cannot be captured from photos alone.
A 2021 paper with code (I haven't tried it): https://github.com/akashsengupta1997/hierarchicalprobabilist... to generate a 3d morph-able human-shape in the correct pose from a single picture.
A picture always has scale uncertainty (a 2-meter human viewed from 1 meter away look the same that a 1-meter human viewed from 0.5 meter away), so that is an additional problem that must be taken care of.
But now recent phone have 3d sensor that provide information that could be useful.
To generate 3d human models there is also makehuman . In the old days there was a soft called facegen, that could generate 3d face models and could automatically fit their parameters to two pictures using an iterating refinement procedure.
Deep-learning usually estimate everything jointly in a single step so they are faster, but often less accurate. But there exist models that learn to refine a previously generated model, so you can apply them repeatedly and get improved quality (Denoising Diffusion Probabilistic Models is one generic class of models that does this).
https://developer.apple.com/augmented-reality/object-capture...
Edit: after looking this is what I was thinking of.
usdz.app
Matrox had a famous 2D to 3D demo
The opposite of this would be causal reasoning but that is hard for humans as well. Not everyone can use causal reasoning, it requires many years of training. The discovery of causal mechanisms (science) is slow and it takes many people to do it.
For example the president of Turkey believes interest rates should be lowered in a hyperinflation. What can we do? Causal reasoning is hard.
I think we do something else, slightly different than causal reasoning. We're generating hypothesis and testing them out, in an iterative process. Like a GAN (generator+discriminator) or Actor Critic (policy + value function). More famously, AlphaGo was generating many rollouts for each move, one model to generate moves, another to evaluate their consequences.