DUSt3R: Geometric 3D Vision Made Easy
dust3r.europe.naverlabs.com
dust3r.europe.naverlabs.com
Getting 3d view from few pictures of an apartment's listing https://twitter.com/JeromeRevaud/status/1764035510236758096
Two pictures of kitchen https://x.com/janusch_patas/status/1764025964915302400
Two pictures of office without any overlap https://x.com/JeromeRevaud/status/1763495315389165963
That's why it can still reconstruct a scene even if the images do not overlap at all: https://twitter.com/JeromeRevaud/status/1763495315389165963
But what that also means is that this is closer to generative AI than to objective measurements. If the image to depth estimation goes very wrong, it might hallucinate shapes that aren't there.
But people do that all the time too. Relying on priors is fine for many practical applications and sometimes there's no way around it.
Too many times I’ve read claims without source so no one can reproduce and verify results. Now I can, and have, verified the results. Top notch.
PYTORCH_ENABLE_MPS_FALLBACK=1 python3.10 demo.py --weights checkpoints/DUSt3R_ViTLarge_BaseDecoder_512_dpt.pth --device 'mps'
Seems like we get more and more generalist approaches which are less specific and combine a lot of what used to be individual steps and techniques. In doing so they don't only become conceptually simpler but surprisingly more accurate as well. Possibly because a unified approach is more integrated and thus better at filling the gaps in one sub-problem with information form other sub-problems.
The fiddly, brittle and multi-step nature of 3D vision endured longer but is going through the same transformation.
> One thing that should be learned from the bitter lesson is the great power of general purpose methods
For example, I tried this with a dog (walking around the seated dog, taking photos as I did so). The dog turned her head while I was taking photos. The portion of the head that moved was not represented in the final output.