Zero-1-to-3: Zero-shot One Image to 3D Object
zero123.cs.columbia.edu
zero123.cs.columbia.edu
If it runs fast enough I wonder whether one could just drive around with a webcam and generate these 3d models on the fly and even import them into a sort of GTA type simulation/game engine in real time. (To generate a novel view, Zero-1-to-3 takes only 2 seconds on an RTX A6000 GPU)
This research is based on work partially supported by:
- Toyota Research Institute
- DARPA MCS program under Federal Agreement No. N660011924032
- NSF NRI Award #1925157
Oh, huh. Interesting. Future Work
From objects to scenes:
Generalization to scenes with complex backgrounds remains an important challenge for our method.
From scenes to videos:
Being able to reason about geometry of dynamic scenes from a single view would open novel research directions --
such as understanding occlusions and dynamic object manipulation.
A few approaches for diffusion-based video generation have been proposed recently and extending them to 3D would be key to opening up these opportunities.With the fairly shallow slope of the GPU performance curve overtime, I don’t see them just Moores Lawing out of it either. this would need two, maybe three orders of magnitude more performance.
~20ms is that threshold, but even 40ms latency is barely noticeable for single player games.
At a 20ms total round trip, that only buys you about a 1500 mile radius, again completely ignoring all other latencies.
For casual gamers and turn based games maybe it could work, as a niche. For FPS, multiplayer, ARPG, and so on, it's a dealbreaker, anything over 100ms feels too sluggish.
We should be happy we have so much autonomy with our own hardware, I don't want some big cloud company to be able to tell me what I can play and render, unless we want the "you will own nothing and be happy" meme to become reality.
On Sonic fiber internet in San Francisco, I get 1.5ms to the POP. It is only 4.5ms to my VM in Hurricane Electrics Fremont DC.
If you look at a graph, that stopped being true well over a decade ago.
You could always get as much parallelism as you wanted by adding more chips.
Furthermore once you've identified the make and model of the car, its relative position in 3d, any anomalies -- that ain't just a Ford pickup, it is loaded with cargo that overhangs in a particular way -- its velocity, 'etc -- I'm quite sure that extrapolating additional information from the subsequent frames will be significantly cheaper as you don't have to generate a 3d model from scratch each time.
I think this is a viable exploratory path forward.
Make it work <- you are here
Make it work correctly
Make it work fast
Edit: Scotty does know ;) Make it work <- you are here
Make it work correctly
Make it work fastNot true.
In the last week, a lot of the ideas I’ve read about in the comments of HN, have then shown up as full blown projects in the front page.
As if people are building at an insane speed from idea to launch/release.
It's definitely a fun time to be involved.
Its outputs can provide the inputs for NeRF training, which is why they mention NeRFs. But it's not NeRF technology.
From the SJC paper:
> We introduce a method that converts a pretrained 2D diffusion generative model on images into a 3D generative model of radiance fields, without requiring access to any 3D data. The key insight is to interpret diffusion models as function f with parameters θ, i.e., x = f (θ). Applying the chain rule through the Jacobian ∂x/∂θ converts a gradient on image x into a gradient on the parameter θ.
> Our method uses differentiable rendering to aggregate 2D image gradients over multiple viewpoints into a 3D asset gradient, and lifts a generative model from 2D to 3D. We parameterize a 3D asset θ as a radiance field stored on voxels and choose f to be the volume rendering function.
Interpretation: they take multiple input views, then optimize parameters (a voxel grid in this case) to a differentiable renderer (the volume rendering function for voxels) such that they can reproduce the input views.
[0]: https://pals.ttic.edu/p/score-jacobian-chaining [1]: https://github.com/awesome-NeRF/awesome-NeRF
Until the oceans boil...
And I'm not sure if it's technically possible for one AI to train another AI with the same algorithm and have better performance. Although I could be wrong about any and everything. :-)
All you have left to do is to AI the process of training AI, kind of like building a lathe by hand makes a so-so lathe but that so-so lathe can then be used to build a better and more accurate lathe.
All of that modern machinery was essentially bootstrapped off a couple of relatively flat rocks. Its going to be interesting to see where this LLM stuff goes when the feedback loop is this quick and so much brainpower is focused on it.
One of my sneaky suspicions is that Facebook/Google/Amazon/Microsoft/etc would have been better off keeping employees on the books if for no other reason than keeping thousands of skilled developers occupied, rather than cutting loose thousands of people during a time of rapid technological progress who now have an axe to grind.
There are tricks to do it faster but they all involve using other vision models that themselves are trained for as long.
And the pretrained vision encoder will have at some point been trained to minimize text-visual token cosine similarity on some training set, so it really depends on what exactly that training set had in it.
https://www.youtube.com/watch?v=zPqJUrfKuqs
Does it stabilize, or refine prejudices, or go on a fractal journey of errors over the weight landscape?
They are precomputed, "Note that the demo allows a limited selection of rotation angles quantized by 30 degrees due to limited storage space of the hosting server." but I don't think they are curated, the seeds probably correspond to the seeds of the live demo you can host (they released the code and the models)
[1] Methods for 3D digitization of Cultural Heritage: http://www.ipet.gr/~akoutsou/docs/M3DD.pdf
How many people have taken you up on that offer? Unless it's a shitty/low-effort painting, it seems insane to me that anyone would destroy their artwork in exchange for an NFT of that same artwork.
It is also important to highlight that we are doing this project at our own risk, with our own money, have built the hardware and software, and not charging artists for the process. Just the primary market sell is split between 85% for artists and the rest for the project. Pretty generous in this risky market.
Please also include the number of those people who actually understand what an NFT is. As a native Miamian, I can guarantee you not a single one does. This city has always been a magnet for the get rich quick scheme types, and crypto is a good match for that because it's harder for a layman to grasp the scam part.
This is what my startup is getting into. So I'm very interested.
These aren't "game ready" - the sculpts are pretty gross. But we're clearly getting somewhere. It's only going to keep getting better.
I expect we'll be building all new kinds of game engines, render pipelines, and 3D animation tools shortly.
For example, Apple supposedly has put some time into 3d asset building (presumably in support of AR world building content).
Can these inference techniques stack or otherwise help more detailed object data collection?
So maybe someday, but I think it would have to be a project that targets CAD.
They're not building this for games they're building it for autonomous weapons.
Here they list "GT Mesh", "Ours", "Point-E", and "MCC". Does anyone know what technique "GT mesh" refers to? Is it simply the original mesh that generated the source image?
That's something to be aware of, especially when you're using convenience data of unknown quality to evaluate your model – many research datasets scraped off the internet with little curation and labeled in a rush by low-paid workers contain a lot of SEO garbage and labeling errors.
Anyone have any contacts? They seem to be extremely elusive
https://www.hollywoodreporter.com/wp-content/uploads/2017/07...
Probably needs a ways to get there but to be able to do robust SLAM etc. With just a single camera would make things much less expensive.
It might be possible to create a "original view + new angle" conditioned model much more easily by taking the Controlnet/T2IAdapter/GLIDE route where you freeze the original model.
Text to 3d seems almost close to being solved.
It also makes me think a "original character image + new pose" conditioned model would also work quite well.