Stable Diffusion 2 Depth Guided model: architecture photos from dollhouse
twitter.com
twitter.com
I already incorporated this to Houdini to generate 3D distance meshes and it works wonders. If anyone is interested in experimenting with this, I really recommend Hugging Face diffusers library.
I wonder when we will be able to generate studio quality 3D models from a single 2D image. I know there are solutions that can do this but it is still nowhere near professional quality.
It’ll get you 80% of the way there for each texture really fast. Would work really well if you are making a ton of NPCs and made a head model and a separate clothes model.
To solve the issue with hidden faces (right now it just projects out copies of the visible textures, you could combine this with inpainting: Rotate the scene / model to expose untextured faces sides that get masked out and inpainted, then rinse and repeat until the model is fully textured.
Another interesting (albeit much harder) thing to try would be to take the image output from SD, run it through MiDaS to get a new depth map with additional details. Diff the depth maps in 3d space and update the geometry to match before projecting the image. Combine it with the first suggestion and you would have a process for progressively refining a detailed, fully textured model.
I made a go at this exact approach last night. I'm new to working with 3d data but at least got a mesh rendered from the depth map. Spent most of my time fighting with tooling. Let me know if you want to help out.
The depth map encodes an array of relative positions, but without knowing the camera settings, field of view, etc. mapping them back to a coordinate is a guess. Luckily if you are rendering your initial depth map from a 3d model, you can use the camera settings to get the information you need to convert a depth map back into a correct array of 3d pixel coordinates.
You can use those pixel coordinates, which you know lie along existing surfaces, to subdivide the mesh to add complexity where needed. Then when you generate the new image and corresponding depth map, you can project those into 3d space using the same camera settings. You will need to filter out coordinates that are past some margin of error from the existing mesh and then use those to push and pull the existing mesh faces.
Map your textures onto your updated mesh and then also generate a confidence texture. The more each face points towards the camera the more confident you can be about the texture there.
Then move / rotate the camera and repeat, but this time also render the model with the confidence texture to generate an inpainting mask image. Continue from different positions / angles until the scene or model has been full textured with a high degree of confidence across the entire texture.
* Generate depth image from 3D meshes
* Send depth image to local HTTP server running Hugging Face diffusers library. Generate textured image.
* Cast textured image to a 2D plane surface that covers the camera extends.
* Traverse the UV surface of the 3D meshes and find closest point on the 2D stable diffusion image plane. Copy the RGB value to UV texture and you’re done.
Thank you for sharing, this is just incredible
The open sourcing of these models was of course instrumental to this. So thanks Huggingface et al., I guess!
The models are openly available but the license places some restrictions on the use.
> You agree not to use the Model or Derivatives of the Model:
> - In any way that violates any applicable national, federal, state, local or international law or regulation;
> - For the purpose of exploiting, harming or attempting to exploit or harm minors in any way;
> - To generate or disseminate verifiably false information and/or content with the purpose of harming others;
> - To generate or disseminate personal identifiable information that can be used to harm an individual;
> - To defame, disparage or otherwise harass others;
> - For fully automated decision making that adversely impacts an individual’s legal rights or otherwise creates or modifies a binding, enforceable obligation;
> - For any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics;
> - To exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm;
> - For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories;
> - To provide medical advice and medical results interpretation;
> - To generate or disseminate information for the purpose to be used for administration of justice, law enforcement, immigration or asylum processes, such as predicting an individual will commit fraud/crime commitment (e.g. by text profiling, drawing causal relationships between assertions made in documents, indiscriminate and arbitrarily-targeted use).
https://github.com/CompVis/stable-diffusion/blob/main/LICENS...
For many, this distinction is more of either an academic one or one where we are OK with those kind of restrictions. Where I live if I did many of these things I would be either criminally or civilly (the right phrasing?) liable so having a license that tells me I can't break the law is a little redundant. I think 4 is possibly legal but an edge case and the last one is nation state level.
You are probably technically correct, but the context here is very important IMO.
7. Updates and Runtime Restrictions. To the maximum extent permitted by law, Licensor reserves the right to restrict (remotely or otherwise) usage of the Model in violation of this License, update the Model through electronic means, or modify the Output of the Model based on updates. *You shall undertake reasonable efforts to use the latest version of the Model.*
I think this means it is now illegal to use SD 1 if at all reasonably possible to use SD 2?
It does add risks for a company sure.
edit - thanks for adding that in, it's an important part of the picture.
This way if a court rules that generated models don't qualify for copyright they can still enforce it under contract law.
Later this week, Draw Things will have a release with depth2img model as well, initially only supporting iOS-provided depth map, but will do MiDaS inferred depth map soon. Going to be exciting to see what people comes up with.
https://github.com/backnotprop/Colab-Stable-Diffusion-2-Dept...
Take a look at Analog Diffusion outputs, for example.
I certainly do not.
Anyway, none of the existing image-gen models can be used in content production in their vanilla state. Prompt to image is fine as a hobby, but style/concept transfer is what makes them usable in professional setting. Or rather will make in nearest future, as this is all still highly experimental. SD in particular is quite small and is not a ready to use product, not intended for direct usage. It's a middleware model to build products upon. Such as Midjourney.
If you're not doing portrait photography, replace Sigma 85mm f/1.4 with something else for example Sigma 24mm for more wide-angle photography.
E.g. 85mm vs 24mm make no specific changes to a photo. SD appears to just interpret these as "make a photo look realistic", and any changes to the photo as you switch between them are simply incidental.
Dreambooth trained on the ground truth photos might help a bit
Photographers have been attempting to emulate various films in different ways forever. Stable diffusion could do it absolutely perfectly.