Magic123: One Image to High-Quality 3D Object Generation
guochengqian.github.io
guochengqian.github.io
Research is additive not a zero sum thing.
lmao are you for real?
From what I understand, they use the text keywords detected from the image as guidance and they also apply a loss between the current diffusion state and the source image. In effect, this is stable diffusion for 3D shapes but with clever conditioning. That means this algorithm will also work just fine if you have 2+ input images.
The more assumptions you can bake into the parameters of some model, the more degrees of freedom you get in the actual measurement process (e.g. reducing the amount of actual data necessary).
But yes, the images totally came from a controlled environment. They rented like 50 similar cameras and hardware-synchronized the shutters.
What would really make a difference if you could make a decent model from 30 images snapped by a drunk teenager on their $200 budget phone in bad lighting.
If money and resources are not an object then yeah it's easy but for most people it is.
The best camera is the one you have with you, and the same goes for this too.
It's not there yet for AAA games, but will get there in time. In the meantime, right now as it is, it's could save days of time for the indie game dev.
Seems valuable to be able to generate 3d models from single 2d illustrations, for games and other media.
And there are probably heaps of other applications as well.
This is just for creative/artistic use, because you obviously don't want software to "guess" on something used for engineering. In that case, you would use an appropriate 3D scan.
Also, 3D assets aren't generally created as scenes. A 3D-2D-3D pipeline also has wide application in AR, VR, standalone modeling, animation, etc.
The problem with the multi-perspective approaches are: to do it well you need lasers and extremely stable AND perfectly localized observation points, making it slow, expensive and fragile for robotic applications, OR the image-to-model component of it is a grotesque anti-physics artifact hallucination which reinforces entropy as much as it does valid data leading to shitty models.
Essentially, this single image to 3D step is the key to both forms.
To effectively mirror this process in machines, we need models that can extrapolate a 3D environment and it's objects from sensors. If we can get an approximation of what something looks like from a different perspective, it's valuable for naviatigation and objective planning.
We could also theorize about the effects it will have on machine's awareness and attention.
I'd like to see how it performs on a novel object that is similar to one in the training set, but not included in the training data.
Gotta say, tools like these or NeRF will only revolutionise 3D modelling if they ever get topology right. That's the hard part.
I think it's equally likely that we'll end up with replacing meshes - or with a hybrid pipeline where non-mesh representations coexist with something else.
I am thinking about realtime mainly but I think the same thing might apply to "offline" rendering (does anyone still call it that?)
Not really. Depending on the use case and adequate tools it can be much faster to the alternative of making these manually as meshes and textures.
Using the banana as an example, if this can be converted to a volumetric model (voxels) the puffied side can be shaved off using sculpting tools and the model be converted to a mesh much faster than making it from scratch. While the end result wouldn't be good for looking at it up front, it can be perfectly viable for background props in a game, especially something that is viewed from a bird's eye view or a drawn out third person perspective (though even up front it'll look better than what you see in some games[0] - and that is AAA).
In fact there have been several games using photogrammetry already to construct 3D models out of taking photos of real places from various angles and converting them to point clouds and then to meshes - which after that they need to be cleaned up by artists. This all takes time, is costly and needs specialized hardware and software and yet developers do it. The linked paper is about a method that significantly lowers those barriers while giving decent results even if they still need to be edited.
There needs to be more foundational work in this field that can outperform or even improve the NeRF-based techniques. And the current herd mentality of researchers should be changed into exploring the alternatives. There is a reason why expensive automobile companies still rely on the physical modelling of their design. It's hard to simulate the physical conditions only through CAD modelling. Sure NeRFs are cool and they can make impressive results. That doesn't necessarily mean it is the means to an end. Look where rasterization brought us! NeRF is like rasterization. It is going to be used. But highly quality graphics was possible through GI and ray tracing! NeRF needs something equivalent that is physically grounded.
Have you seen Plenoxels? https://alexyu.net/plenoxels/
Now, here's a thought to chew on: once we've mastered 2D to 3D, what's stopping us from exploring 3D to 4D conversion? That would take this to an entirely new level.
Next step: a few frames of video to produce a fully textured and articulated 3d model?
It looks like if you use the full pipeline, it takes 2.5 hours for textual inversion, 40 minutes for coarse estimation, and 20 minutes for fine estimation. For one image, on a 32G V100.