SPAD: Spatially Aware Multiview Diffusers
yashkant.github.io
yashkant.github.io
SpeedTree already used billboard textures 10 years ago and that's still the way to go if you need a forest in UE5. Fortnite did slightly improve upon that by having multiple billboard textures that get swapped based on viewing angle, and they call that impostors. But the core issue of how to reduce overdraw and poly count when starting with a high detail object is still unsolved.
That's also the reason, BTW, why UE5's Nanite is used only for mostly solid objects like rocks and statues, but not for trees.
But until this is solved, you always need a technical artist to make a low poly mesh onto whose textures you can bake your high resolution mesh.
That's still ultimately triangle meshes though, not some other weird representation like NERF, or distance fields, or voxels, or any of the other supposed triangle-killers that didn't stick. Triangles are proving very difficult to kill.
This illustration from the page you linked to shows that as well:
https://cdn2.unrealengine.com/nanite-in-fortnite-chapter-4-p...
The alpha masked holes move around, but the polygons remain static. That means if you draw a tree with this, you still have the full overdraw of the highest-poly mesh.
https://cdn2.unrealengine.com/nanite-in-fortnite-chapter-4-t...
https://cdn2.unrealengine.com/nanite-in-fortnite-chapter-4-t...
I suspect this will continue to be an uphill battle because there aren't billions of high quality 3D models and PBR textures just lying around on the internet to slurp up and train a model on, so they're having to build it as a second-order system using images as the training data, and muddle though the rest of the steps to get to a usable 3D model.
I think however, that the existence of such tools could perhaps motivate more people to create low quality 3D assets, that for many purposes are good enough - but people might be shy to do that due to many high quality assets shaming them... Once AI 3D stuff will flood the space, amateurs might start competing with it - or so I hope.
Also it's just a matter of time until we reach the 80% (from the 80/20 rule), 90% and 98% quality of these tools. The remaining 2% will still be of value in AAA titles though...
https://github.com/bytedance/MVDream-threestudio
It should be fairly straightforward to adapt this repo (which is based on threestudio) to use SPAD instead of MVDream.
I.e. can these be assets for traditional game engines ?
Could a sort of photogrammetry like Meshroom or Reality capture do the trick ?
And for in-painting I think you’ll find text-to-image is still useful to artists. It’s extra metadata to guide the generation of a small portion of the final image.
(Also, reference images can absolutely be used to communicate style to a diffusion model)
What are your ideas about the differences between a human and AI's creative process ?
Are there any similarities, or analagous processes ?
Do you think creators have an kind of latent space where different concepts are inspired by multi-modal inputs ( what sparks inspiration ? e.g. sometimes music or a mood inspires a picture ) and then the creators make different versions of their idea by combining different amounts of different concepts ?
I am not being snarky, I am genuinely interested in views comparing human an AI's creative processes.
Most illustration briefs are also not wrote descriptions of images because people are remarkably bad at describing what they want in an image, beyond in the most general sense of its subject. This is why you see DALLE doing all kinds of prompt elaboration on user inputs to generate “good” images. Typically, the illustrator is given the work to be illustrated (e.g. an editorial), distills key concepts from the work and translates these into various visual analogues, such as archetypes, metaphors and themes. Depending on the subject, one may have to include reference images or other work in a particular style, if the client has something specific in mind.
We’re building a model optimized for the machine, not people.
Artists can go collect clay to sculpt and flowers to convert to paint. Computers are their own context and should not be romantically anthropomorphized
In the same way fewer and fewer people go to church, fewer and fewer will see the nostalgia in being a data entry worker all day. Society didn’t stop when we all got our first beige box.
No one ever set a goal for AI to achieve the same result; just replace labor.
Your post is the same dull strawman non-engineers (I also have a BSc in engineering and MSc in math) repeat about AI.
Find me a formal proof of how “images are made” and I’ll show you one possible model of an infinite number of possible models to explain it with a few axiomatic correct twists to the math since all of our symbolic logic is a leaky abstraction that fails to capture how anything is “fundamentally made”.
Pretentious semantic wank is all you’re shipping
With a local model you can find latent space coordinates any way you want and patch the pixel generation model any way you want too. (the above are usually called textual inversion and LoRAs.)
I would personally like to see a system that can input and output layers instead of a single combined image.
Yay ambiguous acronyms.