A working implementation of text-to-3D DreamFusion, powered by Stable Diffusion
github.com
github.com
There is also an alternative way to handle this latent difference with the original paper that should also work :
Instead of working in voxel color space, you push the latent to the voxel (Aka instead of having a voxel grid of 3d rgb color, you have a voxel grid of dimlatent latents, (you can also use spherical harmonics if you want as it works just the same in nd) ).
Only the color prediction network differ, the density is kept the same.
The NERF then directly render to the latent space (so there are less rays to render) which mean you need to decode it with the VAE only for visualization purposes and not in the training loop.
Basically if I’m reading it right, this does the synthesis in latent space (which describes the scene rather than rendering vocals) then translates it into a NERF. It sounds kind like the Stable Diffusion description that was on here earlier.
In these, you represent the world as a sparse latent representation, collection of 3D coordinates, and their corresponding SIFT feature descriptor. You rotate and translate these keypoints to obtain a novel view of the features in 2D image. (The descriptor can be taken as an interpolate of the descriptors weighted by the difference in orientation between views as features only match for small viewing angle difference (like 20°) ). And then you could invert features to retrieve a pixel-space image (for example https://openaccess.thecvf.com/content_cvpr_2016/papers/Dosov... ) (although it's never needed in practice.
Coming back to NERF, it's the same principle. When your NERF has converged, if you don't have transparent object, along the ray the density will be 0 except when you intersect the geometry where only a single voxel will be hit in which case you fetch the latent stored in the voxel latent (spherical harmonics) from the direction given by the ray during the training with a latent image.
The rendering equation is still the same but instead of rendering a single ray, it would be analog to rendering a group of close rays in order to render a patch of image, of which the latent is a compressed representation. You have to be careful not to make the patch too big, because like with a lens in the real world, spherical transform flip the patch-image upon translation, but neural network should transparently handle this.
The converged representation is an approximation based on linear approximation and interpolation along positions and ray direction, provided that you have enough resolution, you can construct it manually from the solution and see how it behaves in the rendering.
Will the convergence process work ? It will depend on how well latent mix, along a ray. The light transport equation is usually linear, and latent usually mix well linearly (even more so when weighted by a density), But in the case it doesn't mix well you can learn a mixing of latent rule that help it converge.
Also once you have a latent Nerf, it won't allow you directly to obtain a STL/obj directly but you should have 3d consistent views from which you could render a classical NERF, but you can/should also instead optimize for the classical voxel grid, that fit the latent voxel grid (aka that give the same image patches).
for d in ['front', 'side', 'back', 'side', 'overhead', 'bottom']:
text = f"{ref_text}, {d} view"
https://github.com/ashawkey/stable-dreamfusion/blob/0cb8c0e0...It then sprinkles some noise on the rendering, makes Stable Diffusion improve it a little, then adjusts the voxels to produce that image (using differentiable rendering.)
Rinse and repeat for hours.
That's interesting for a couple of reasons. I can see why that works. It also implies that for closed objects, the voxel data on the interior (where no images can see it) will be complete noise, as there's no signal to pick any color or lack of a voxel.
text = f"{ref_text}, front cutaway drawing"
Maybe?It certainly doesn't look as good as the original, yet. I wonder if that's due to the implementation differences noted, less cherry picking in what they show, or inherent differences between Imagen and Stable Diffusion.
Maybe Imagen just has a much better grasp of how images translate to actual objects, where Stable Diffusion is stuck more on the 2d image plane.
Non-Euclidean back-polygon imaging? Good work, algorithm. ;)
Right now, we are relying on the Sketchfab API to populate our (Blender) scenes, which is an imperfect lens through which to visualize the contents of texts that our non-technical "clientele" are studying.
Since we are publishing these scenes via WebXR (Hubs), we have specific criteria related to poly counts (latency, bandwidth, etc) and usability. Regarding the latter concern, it's not clear that our end users will want to wait/pay for compute.
*copyedited
Like this (nsfw content): https://lexica.art/?q=Intricate+goddess
There is an addictive and trippy quality to this and it is yet to hit mainstream -- The art itself is stunning but it goes beyond that, the ability to nudge it around and make variations to it is incredible. now add the fact that you can train it with your own content. people are going to go bonkers with this and it's going to open up a lot of debates too.
# test (exporting 360 video, and an obj mesh with png texture)
python main_nerf.py --text "a hamburger" --workspace trial -O --test
So I guess so. That's pretty awesome.I just got my 3D printer and was a bit too tipsy to assemble it the day it arrived - and have several things I want to print…
It will be interesting to experiment with describing the thing I want to print with text instead of designing it in SolidEdge and see what AI thinks….
I wonder if you can feed it specific dimensions?
“A holder for a power supply for an e bike with two mounting holes 120mm apart with a carry capacity that is 5 inches long and 1.5 inches deep”
Maybe two papers down the line :D For now you might have more luck with something less specific.
I like my DALLE expressions of “masterchief as ventruvian man as drawn by da Vinci”
And my “technical exploded diagrams of cybernetic eco skeleton suits in blueprint”
Try those out?
At the beginning of this year, most technical people would have told you that graphic design was a decade from being automated, and creative video production more.
Now we are at "months and months".
I do feel that the new-gen mechanical CAD will be based on a dialog between a human and a generational suite. Whether it's based on diffusion or other method - that remains to be seen.
Combine that with an AI that can generate an initial design based on a rough specification, have the user apply some constraints and then iteratively have the AI generate new designs, based on known good designs, that fulfil those constraints could be very powerful.
That makes this Stable-Dreamfusion adaptation even more promising.
But I guess what creative industry means to you? Pumping out web UIs or 3d gaming models were never, for the most part, the creative industry; learning to see what people like and copying that for different situations is not necessarily creative and thus what AI easily does; anything that doesn’t come with a lot of learning and practice and talent outside manual work will be replaced by AI soon; the other stuff will take somewhat longer.
If you think this can replace you, you weren’t/aren’t in the creative industry. Same goes for coders afraid of no code.
Edit; but you are also implying you think your job is gone with stuff like this? What do you do? Also I am hoping I will be replaced: I have been thinking I will be replaced since the early 80s as my work as a programmer is not so exciting (I love it and will keep doing it even if it’s not viable anymore, which I do believe for the 20% of people who do niche work is very far off AI wise; like I think with creative as well) but it seems closer now than ever.
Edit2: looking at your profile work, you don’t seem you will be replaced by anything soon; what is the anger about? Do you have public blogs/tweets about your feelings about this; looking at your work (in your HN profile) you seem the group not touched by this at all.
Politically, keeping the gains of this not flowing to a few companies or individuals should be a priority; as has been proven (and been said by Carmack), AI is trivial after invention and will be cloned by open source in days/weeks; then trained sets is key to making sure we all benefit. This needs to remain the case to make sure society works imho.
But yeah, soon is 10 years for many things I feel now and 100+ (infinite) years for others. I observe that things that are copilot for me now, are things that millions of programmers cannot even come up with and those need to do something else. And that’s close.
The talented people is further away because of ‘understanding’; statistical pattern matching is maybe (we don’t know) something different than understanding; when we manage to use temporal flow and so, conversational prompts, which a lot of people are researching and developing now, it will get really interesting. So far it seems inevitable… until the next AI winter.
What I am concerned and angry about is the next generation. The people who are just becoming pros, who are at the point where they are glad to get a job cranking out hundreds of models of sneakers for EA's Basketball Jam 2024 or doing a bunch of D&D character commissions or whatever other commercial art job because it is paying the bills (including working off their massive student debt) by doing what they actually trained to do instead of some shitty minimum-wage job or an even shittier "gig economy" thing. That's the window that all this AI art crap is making a lot smaller.
An AI winter might happen though if we don’t move on from this place.
One new law that explicitly redefines "fair use" to exclude "scraping half the entire internet and dumping everything you find into a training dataset", and creates a new framework for proper licensing of training data along with hefty penalties for distributing datasets without these licensing would be a huge roadblock.
A grassroots effort of pissed-off artists starting a sideline in assassinating AI researchers would have a pretty chilling effect on the field, too. This may be a little extreme. There are probably solutions that don't go this far. It sure does make a pleasant revenge fantasy though. Time to go read some accounts of how the Unabomber was caught...
Automation creates jobs rather than destroying them. What destroys jobs is mainly bad macroeconomic conditions.