DreamFusion: Text-to-3D using 2D Diffusion
dreamfusion3d.github.io
dreamfusion3d.github.io
From a pile of random undifferentiated images the model has learned the detailed 3D structure and plausible poses and variants of thousands (millions?) of everyday objects. And all we needed to get that 3D information out of the model was the right sampling procedure.
The example with a squirrel wearing a hoodie demonstrates an interesting edge case, the "front" of the squirrel (with hoodie over the head) show a normal hooded face as expected, but when you rotate to the "back" you get another face where the hoodie is low over the eyes. Each looks fine in isolation, but in aggregate it seems like we have a two-faced squirrel.
It'll just make up some colours and geometries that don't contradict anything it already knows from the defined perspectives.
Or leave it empty.
Bad prompt, missing implied antecedent/ambiguous subject...
You may want:
Back view of a cat which is wearing sunglasses, back view of a cat, but the view is wearing sunglasses, etc... I actually tried using projective terms from drafting books, and didn't get great results. Nor anatomicals either.
In short: natural language is not good enough and you need a DSL. If only the last 60 years of language research had warned us of this.
Up next: English sentences are ambiguous and need context information to parse correctly. Machine learning community in shambles.
The trouble is all those darn uninitiated and trying to create a generalized oracle to map their inspecific ramblings to what they mean to free them of having to actually communicate properly...
Actually, funnily enough, this has cross section with philosophy in a way most programmers scoff at; but communication is frigging hard, and worse, detecting when someone is trying to get something across, but just needs a nudge in the right direction to be able to find the language to explain it is really damn hard.
I run into it every time I get a haircut. I have no idea how to speak their language, so it's always "Uh... A little off the top and rounded at the back, I guess?"
-the point
But that's not really surprising because when you have enough data, even simple clustering methods group objects like faces by the direction they are looking to. With enough views even a simple L2 distance in pixel space allow t-SNE to do that.
They are injecting the 3D constraints via the NERF and an optimization process to add the consistency between the frames.
It's a deep dream process that optimize by alternating updates for 3D consistency, and updates for text-to-2Dimage correspondence. It's searching for a solution that satisfy these two constraints at the same time.
Even though they only need to run a single diffusion step to get the update direction, this optimization process is quite long : 1h30 (but they are not using things like instant Nerf (or even simple voxel grids) ).
But this will allow for creation of a dataset of 3D objects with corresponding text, which will then allow to train a diffusion model that will have a 3D understanding and will be able to generate 3D objects directly with a single diffusion process.
Humans don't get natively 3D training data either. We don't have depth sensors or any other natively 3D senses. Our eyes only ever see 2D images and we learn 3D structure from that, so I guess these 2D image models are doing something analagous. Except they don't even have the benefit of stereo images!
I think that would be an oversimplification. We do have some 3D information from focus and eye convergence.
If we elevate "correlations" to mean "understanding", we're quickly going to run out of words.
At some point we have to accept that when you layer simple ops on top of simple ops enough times you get complex behavior.
And the submarine clearly can swim! :D
The model encompasses 3D information, but a model is not a thinking entity able to understand anything. Well, it might be argued that saying that the "model understands" is a metaphor, and used as such OK, fine. But that’s it. Analogies rarely scale well.
As for humans, have you ever heard of proprioception? If that is not "native 3D sensors", I have no idea of what a 3D sense might be.
A machine with 3d understanding from 2d samples predates computers
This seems like basically plugging a couple of techniques together that already existed, allowing to turn 2D text-to-image into 3D text-to-image.
They have others as well.
as with a majority of ML research
Plus "we did the same thing, but with 10x the compute resources".
But yeah.
Do this enough times and eventually the thing you have looks indistinguishable from something completely novel.
In his Lex Fridman interview, John Carmack makes similar assertions about this prospect for AGI: That it will likely be the clever combination of existing primitives (plus maybe a couple novel new ones) that make the first AGI feasible in just a couple thousand lines of code.
I also liked the analogy he made with his earlier work on 2D and 3D graphics engines where taking a few short cuts basically got him on a path to success. For a while we had this "almost" 3D capability long before the hardware was ready to do 3D properly. It's the same with AGIs. A few short cuts will get us AI that is pretty decent and can do some impressive things already - as witnessed by the recent improvements in image generation. It's not a general AI but it has enough intelligence that it can still do photo realistic images that make sense to us. There's a lot of that happening right now and just scaling that up is going to be interesting by itself.
I think finding a niche for humans has proved difficult especially because of those reasons, and AI can take those hurdles much easier.
It takes nature thousands of years to create a rock that looks like a face, just by using geology. A human can do that in a couple hours. And then this AI can generate 50 3d human faces per second (assuming enough CPU).
It could be that an AGI is around the corner, as they say. We might not be machines, but are way faster than nature at reaching places. We don't have the option of waiting for thousands of years.
and I agree with you!
And the OP comment its by the magnanimous/infamous AnigBrowl
You need to start doing AI legal admin ( I dont have the terms, but you may - we need legal language to control how we deal with AI)
and @dang - kill the gosh darn "posting too fast" thing
Jiminey Crickets I have talked to you abt this so many times...
I'm a kotlin programmer so it's a scary fun thought for me.
This should be further up than all the speculation about AI accelerationism. There's a very simple explanation why a lot of awesome papers come out right now, it's prestigious conference paper deadlines.
WTF - the singularity is closer than we thought!!!
Whatever it is, this is a massive phase shift.
The pace at which methods scale up is currently a lot faster than hardware improvements, so unless these scaled up methods become incredibly lucrative (not impossible), I think it's quite likely we'll soon-ish (a couple years from now) see a slowdown.
The numbers of researchers and research labs scaled up that there is now many well funded teams with experience.
Public tooling and collaboration has reached a point where research happen across the open internet between researchers at a pace that wasn't before possible. Common Crawl, stable diffusion, hugging-face, etc...)
All the techniques that took years in small labs to prove as viable are now getting scaled up across data and people in front of our eyes.
I'm curious where this will end up in a year. Will it plateau? If so, when?
Although they start with random initialization and a text prompt. It seems to work well. I now see no reason we can't start with image initialization!
> In the days when Sussman was a novice, Minsky once came to him as he sat hacking at the PDP-6.
> “What are you doing?”, asked Minsky.
> “I am training a randomly wired neural net to play Tic-Tac-Toe” Sussman replied.
> “Why is the net wired randomly?”, asked Minsky.
> “I do not want it to have any preconceptions of how to play”, Sussman said.
> Minsky then shut his eyes.
> “Why do you close your eyes?”, Sussman asked his teacher.
> “So that the room will be empty.”
> At that moment, Sussman was enlightened.
It’s totally sci-fi, and at the same time seems to be possible? I am amazed how even image generation evolved over the last year, but that’s just me daydreaming.
How long then until we get photorealistic AI generated 3D games and experiences?
But I agree that particular company is creating a bad marketing around the term Metaverse. (But pretty decent HW.) NVIDIA has much better footing with their Omniverse. But in the end, we all know the thing will be build as web browsers are today. USD + JS + WebRTC + WebXR can go a long way.
- they rarely provide the data or code used so it's basically "i swear it works bro" research
- what they achieve is usually through having the most pristine dataset on the planet and is often unusable by other researchers
- other times they publish papers that are basically "we slightly modified this excellent open source paper, slapped an internal name on it and trained it on our proprietary dataset"
- sometimes they achieve remarkably little but their papers still get a shiny spot because they're a big name and sponsor all the conferences
- they've also been caught trying to patent/copyright ML techniques; disregarding that this is the same as privatizing math, these are often techniques they plainly didn't come up with
Also ever since OpenAI did their "we have to go closed-source for-profit to save humanity" PR campaign, every company that releases models that can achieve a large amount in NLP/CV gets dragged by the media and equated to Skynet.
Also it’s one of the only papers they put out that falls completely outside what I put above. They released everything about it including the model code, pretrained weights, the techniques; and it took quite a while for the model to “catch on” while it was peer reviewed and reproduced by others.
Something something broken clock
Once the paper is accepted (or rejected) the names may be revealed.
Though, in reality, the reviewers can often easily tell who wrote the paper.
https://dreamfusion-cdn.ajayj.com/gallery_sept28/crf20/a_DSL...
But if you look closely, the pin looks like it's actually rolling across the dough as the camera orbits.
where-by learning that Pixar was developed by steve jobs when lucas didnt think there was a future for computer animation... and so steve bought the death star from lucas...
That became pixar...
AI is going to fucking kill it - what will happen in the next decade will be ANYONE uploading a script to an AI to make a full length movie...
AND their will be editing tools as well that are AI driven...
Like mentioned by William Gibson
*The future is here, its just not evenly distributed yet*
All the primitive components for this future seem to be at an early stage of inception. We can't say for sure whether they will mature to the point where they can replace us, but the trajectory is certainly pointing in that direction.
So what we will get in coming years / decades - the timeline is anyones guess - is movies acted out to camera in a rehearsal space, and that combined with a(n AI augmented) script to generatively create a film / character driven interactive game).
Nah. These techniques will definitelly lower the barrier for making stuff (not just movies), but that has been case will all transformative technologies.
Before computers, if you wanted to shoot and edit a movie, it was a challenge. Now, you can shoot a movie with your pocket computer and edit it on the same device while shitting. Upload it to video sharing web and billions of people can watch it.
This class of technologies will enable creative people to make a lot of stuff, to iterate quickly. But don’t be naïve that everybody will do that. 99% of that will be trash, and that is fine. I like that it will enable individuals to bring their visions into this world without any need for collaboration. And when highly artistic individuals will begin to collaborate using these tools, that will be an awesome inflection point for art as we know it.
Blown away by how quickly this stuff is advancing, even as someone who's relatively cynical about AI art.
I'm curious about the process you and the team uses to do this. Additionally there's the meme that these things are appearing every hour now, so it could be good for some perspective like "well actually it took n weeks"
Research progress is unpredictable and these advances are not inevitable. We have the privilege to work in an environment full of amazing colleagues and powerful models, but at the end of the day it took a persistent team and a bit of luck.
It trivializes it, in my opinion.
When asked the question of is lambda/GPT-3 and/or DreamFusion and it's derivatives an aspect of sentience? there's always a bunch of people who are repeating the same cliche negative line, of "no, it's only attempting to statistically mimic sentience." I agree with the reasoning.
But have we considered the other side of the story? That yes, the mimicry is All sentience actually is. Nothing more.
For example just 2 weeks ago there were two other gaps that you could've used in your example. You could've said 3D interpretation of images wasn't possible and the creation of animated movies wasn't possible and you could've said these few things suggest that there's a gap in what the human brain is doing and what the AI is doing and just throwing more data at the problem doesn't fix it.
Those two examples would be irrelevant today as both of those gaps have Effectively been crossed.
See what I'm saying here. There's two ways of looking at it even from your perspective... Either that gap is so large that the human brain is completely different. Or the gap is small, trivial and will be crossed very very soon.
It suggests to me that what most current approaches are doing is something fairly different from what human / animal intelligence is doing in important ways. That means we will likely continue to see AI do increasingly amazing things while at the same time struggling to perform a lot of tasks that are quite basic for humans.
It is the fact that AI is proving to be a better artist than most humans while not being able to do many things that are simple for a 4 year old that suggests strongly to me that some of the fundamental mechanisms are fairly different still, or current AI approaches are missing some key insights.
I could be wrong. I'd bet money that I'm right if there was an easy way to do it though.
If someone asked this question when GPT-3 came out there'd be thousands of negative retorts throwing out the same tired lines. Now this statement is getting harder and harder to refute.
Stable diffusion has been made available to the public for quite a while now and if anything has disproved a lot of the ungrounded nonsense that made companies like OpenAI censor their generative models.
What's with the STEM reference? Are you implying that STEM is related to intelligence and that people without a STEM background are not intelligent?
It is well known among academics (AKA STEM MAJORS) that human society is a chaotic system and that ML can change society for the better or for the worse... the outcome is basically unknown... It is therefore the intelligent choice to consider the negative consequences of this technology.
To not consider the other side indicates a lack of something.
But we are one month into the experiment, and the technology still struggles to get wholly authentic looking media. I'd say it would be wise to give it a few years before claiming victory.
I watched this humor group when I was little (think 30 years ago). Usually they are inoffensive, but they have this particular sketch were a (man crossdressed as a) woman complains about "his husband hitting her" ... and you can see that one of her eyes is black. With canned laughs during the whole thing.
This used to make people laugh. 30 years ago. Then we had social changes and gradually our perspective as a society shifted, and this humor piece looks ... grotesque.
That's how I interpret the "societal changes" the OP is referring to. We have shiny new things every week and there's no time to adapt to them.
It's used to mean that the placement of the lower level components - vertices, edges and faces - are well aligned to the higher level structure of the object, and make up a well defined 2D grid that flows along the models surface. In particular you'd want edges going along/perpendicular to mesh structures such as limbs, and around facial features and other details in a logical manner.
Otherwise when applying deformations as part of an animation the model will not have the detail in the right places to still look good, e.g. if there is no edges perpendicular to a joint in a limb, the bent version of the limb cannot have a clear smooth line along the joint, and the edges and faces become janky.
Under this definition, marching cubes cannot produce good 2D topology, as the mesh features are all aligned to the cardinal grid instead of the features of the object represented.
You'll have to wait a few months for someone else to replicate it.
I made the same thing the old fashioned way. Mine can actually play though. https://twitter.com/LeapJosh/status/1423052486760411136 :P
The samples are lacking definition, but they're otherwise spatially stable across perspectives.
That's something that's been struggled with for years.
I'm also not sure someone couldn't cleverly figure out a way to use stable diffusion to write code based on a text prompt.
"Mesh exports Our generated NeRF models can be exported to meshes using the marching cubes algorithm for easy integration into 3D renderers or modeling software."
Like they say, "This is the start of something big."
I'm excited and scared. The world is going to look very different in 10 years!
So, authors today often submit anonymously following the conference guidelines, but simultaneously post publicly elsewhere, walking a fine line as not to overstep the conference policies. This appears to be the case for this submission. Note, recently, some conferences, such as CVPR, have started to institute new policies forbidding social media promotion until acceptance, as they adapt to the changing landscape of social media promotion. If this were a CVPR submission, the authors would not be allowed to tweet publicly about their work yet, nor have the version of the webpage with their names visible.
"For he spoke, and it came to be;
he commanded, and it stood firm."
Psalm 33:9, NIV
:)
"transform this into this realistically"
ILM's holy grail