Depth Pro: Sharp monocular metric depth in less than a second
github.com
github.com
I'd bet that if you did that on these examples you'd see that the hair, rather than being attached to the animal, is floating halfway between the animal and the background. Of course, depth mapping is an ill-posed problem. The hair is not completely opaque and the pixels in that region have contributions from both the hair and the background, so the neural net is doing the best it can. To really handle hair correctly you would have to output a list of depths (and colors) per pixel, rather than a single depth, so pixels with contributions from multiple objects could be accurately represented.
They do this. See figure 4 in the paper. Are the results cherry-picked to look good? Probably. But so is everything else.
We plug depth maps produced by Depth Pro, Marigold, Depth Anything v2, and Metric3D v2 into a recent publicly available novel view synthesis system.
We demonstrate results on images from AM-2k. Depth Pro produces sharper and more accurate depth maps, yielding cleaner synthesized views. Depth Anything v2 and Metric3D v2 suffer from misalignment between the input images and estimated depth maps, resulting in foreground pixels bleeding into the background.
Marigold is considerably slower than Depth Pro and produces less accurate boundaries, yielding artifacts in synthesized images.
You're correct about that, but for something like matte/depth-treshold that's exactly what you want to get a smooth and controllable transition within the limited amount of resolution you have. For that use case, especially with the fuzzy hair, it's pretty good.
Each pixel takes a multiple measurements over time of the intensity of reflected light that matches the emission pulse encodings. The result is essentially a vector of intensity over a set of distances.
A low depth resolution example of reflected intensity by time (distance):
i: _ _ ^ _ ^ - _ _ d: 0 1 2 3 4 5 6 7
In the above example, the pixel would exhibit an ambiguity between distances of 2 and 4.
The simplest solution is to select the weighted average or median distance, which results in "flying pixels" or "mixed pixels" for which there are existing efficient techniques for filtration. The bottom line is that for applications like low-latency obstacle detection on a cost-constrained mobile robot, there's some compression of depth information required.
For the sake of inferring a highly realistic model from an image, Neural radiance fields or gaussian splats may best generate the representation that you might be envisioning, where there would be a volumetric representation of material properties like hair. This comes with higher compute costs however and doesn't factor in semantic interpretation of a scene. The Top performing results in photogrammetry have tended to use a combination of less expensive techniques like this one to better handle sparsity of scene coverage, and then refining the a result using more expensive techniques [1].
Plausibly, you can train a model that encodes sufficient information about a specific set of imager+lens combinations such that the lens distortion behavior of images captured by those imagers+lenses provides the necessary information to resolve the scale of objects, but that is a much weaker claim than what monodepth researchers generally make.
Two notable cases where something like monodepth does reliably work are actually ones where considerably more information is present: in animal eyes there is considerable information about focus available, let alone the fact that eyes are nothing like a planar imager; and phase-detection autofocus uses an entirely different set of data (phase offsets via special lenses) than is used by monodepth models (and, arguably, is mostly a relative incremental process rather than something that produces absolute depth).
I'm a little surprised there isn't more research on producing depth from short videos, though, since iPhones take Live Photos and presumably could use more of the natural movement and shakiness to interpret more of the true depth in the scene... Presumably much better than processing a single still.
If you look where a lot of the money has gone into monodepth, self-driving or driver assistance isn't too far away...
Actually, I tend to think self-driving is one of the few places you can make a case for monodepth, as a backup for failures in the rest of your depth sensing suite. You wouldn't want to use it as a primary sensor, but if you're a vehicle driving on a highway and you take critical damage to some of your sensors you do still have to keep driving, if only long enough to get safely off the road, and having something that only needs a single camera is very valuable.
> You wouldn't use monodepth for self-driving
I don’t think you need accurate metric depth for self-driving. “seconds to impact” plus a reasonable estimate of one’s speed (one second from impact at 100km/hour is more dangerous than at 1km/hour both because braking distances grow faster than linear and because a hit at lower speed is less damaging) is way more useful than “meters away” for that.
So, while scale models are indeed a fundamental problem for monodepth, this is not often a problem in practice. The human visual system has the same limitation. A self driving car camera will never in practice find itself peering into a little box of toy cars. And a phone camera will very rarely be looking into a dollhouse instead of a real house. In the latter case, processing can usually be done based on relative depth. The former does need absolute depth, but it might be safe enough to ignore the case where all the cars and surrounding scenery have been shrunk. And if you don’t want to do that, you have wheel odometry with absolute scale you can compare with the imagery.
Yes, and for controlled industrial environments where the domain can be effectively captured in the training dataset, this is plausible. Although a fair number of industrial applications do contain very similar objects at multiple scales, and I don't think much of the existing monodepth work has been meaningfully trained or tested on these cases.
> The human visual system has the same limitation
Only when looking at picture, not when looking directly at the scene even if you only have one eye.
Human vision (and that of many animals, even very small ones like jumping spiders) uses information about the focusing of the eye itself to effectively recover real depth (this mechanism, like structure-from-focus on cameras, is obviously more effective at close distances with shallow depth of field and less effective as the distance and depth of field increase). And, of course, the scanning + focusing behavior of the eye is very different than a perfectly-exposed global-shutter planar imager with deep depth-of-field (which is what the "ideal" car camera would be).
> A self driving car camera will never in practice find itself peering into a little box of toy cars.
Ignoring the case of someone deliberately trying to confuse the car, it will find itself looking at a large variety of children, of different ages, body shapes and sizes, and clothing - and a scale error there can easily have fatal results. Lest this be considered far-fetched, one particular application of monodepth that has been considered by automakers is car backup/360-view cameras, an application area that exists largely because drivers back over small children with alarming frequency.
In real life, you'd use these models for synthetic depth-of-field, adding fake bokeh to a very sharp image that's in focus everywhere. so this seems too easy?
Impressive latency tho.
https://youtu.be/pLfCdI0mjkI?si=8K7rPHu558P-Hf-Z
I assume the first pass is the depth inference here.
Of course, lots of text-to-image models generate a mess, because their training sets are highly contaminated by the messes produced by “Portrait mode”.
I would expect an ML algorithm actually trained to blur the background behind hair to be able to do better (where the resulting image is the output), and maybe a radiance field or splatting algorithm could also do better. In the latter cases, the hair would be separately represented as a foreground object/source.
I could easily be missing something, though. And I did not find any good full-resolution examples.
Adobe actually has a definite ML version in Photoshop under their "Neural Filters" named "Depth Blur (beta)", which is actual AI, but I have slightly poorer results from that than I do with their mainstream "Lens Blur" on Camera Raw.
https://i.pinimg.com/originals/00/f4/8c/00f48c6b443c0ce14b51...
?
Here's the depth map:
(basically, it pretty much thought the whole thing was entirely flat with some distinction of the far-off background)
Gaussian splatting had me amazed, but of course, I'd like a mesh, so there would have to be post-processing to optimize the splats into surfaces, and then reconstruct surfaces.
So this just makes it easier to swap it out without making any other changes.
I'd suspect that a good model will infer what it can from universal depth cues (E.g. Bokeh and perspective) when present but not perform as well if they're as not there or as well as it does on familiar objects.