Moving Camera, Moving People: A Deep Learning Approach to Depth Prediction
ai.googleblog.com
ai.googleblog.com
What do they use to reveal it to the world? GIFs!
Be kind. Don't be snarky. Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
Grayscale disparity/depth maps are somewhat misleading - the large regions of constant intensity suggest that the algorithm is good at segmenting areas of constant depth. However, the flickering in the map suggests that if you actually tried to plot this in 3D, it'd be pretty noisy. Not to disparage the result, but 2D depth/disparity maps tend to look better than what they represent.
You can see this in the synthetic camera wiggle video, focus on the actor's hands, for example.
You can also see this effect in the Stereolabs Zed promo video.
I’d love for OpenSFM or OpenMVS (check github) to get this kind of software.
Also would love to see an implementation of this on github, but hopefully that will follow in time.
Photogrammetry works exceedingly well because the depth maps that they generate are quite precise and accurate, and mesh reconstruction usually assumes that these points are quite close to ground truth.
Deep learning approaches usually have medium accuracy but low precision, which causes the flickering and smooth surfaces that you see on the person. Even the background has flickering despite being computed through stereo, likely because the camera motion is primarily forward-backward (vs. more accurate side-to-side motion), the baseline is likely small, and the depth isn't globally optimized.
This type of research is super great for applications requiring lower accuracy, typically visual-only applications (e.g. selective blurring, faking stereo on a frame, etc.). But as an input to photogrammetry — probably not anytime soon, until the problems above get resolved.
However I do ultimately seek a low accuracy “visually approximate” 3D scene that I could use for simulation purposes. I guess I could rephrase my desire as: I’d love to see this kind of approach used to train an end to end deep learning photogrammetry system. I feel like the parallel nature of neural nets as well as their ability to approximate results could result in a much less computationally intensive solution to my photogrammetry desires.
(I want to train my four wheel drive robot to follow forest trails using the training method described in the “world models” research paper, which requires a simulation to work.)
Full simulation with realistic 3D spaces, enables embodied agents to interact and learn from real-world spaces. Not forest trails, but a real world environment.
If you really want to create a 3D model of forest trails, photogrammetry should be sufficient, because forest scenes are richly-textured.
As far as photogrammetry of forest trails, I found it to be very computationally intensive (taking a GCE 32 core instance 30+ hours using 90+GB of ram to compute a scene, only with errors that made it unusable). It felt very heavy handed and given all the great work I've seen in scene understanding using neural nets, it seems like deep learning would be a promising approach here. Maybe there is commercial photogrammetry software that has better pipelines, but I want to be able to compute my scenes on linux and use hundreds of images.
I did my computation with OpenSFM and OpenMVS. Both wonderful projects for being free and open source. I did get a lot of great results. But I am convinced a simpler way is possible with deep learning.
Also, one of the main steps of mesh reconstruction is depth map generation. It typically takes anywhere from 30-75% of compute time for dense reconstruction, IF it's parallelized thru GPU. If you're using the CPU only to calculate depth maps, you're probably slowing yourself down by an order of magnitude.
If you have a GPU, and use a better SFM-MVS solution, then you can quite easily reconstruct datasets of 1k-10k images within 24 hours.
6d.ai uses depthnets in its mobile photogrammetry pipeline. demo: https://twitter.com/mattmiesnieks/status/1106722396889702406
I'm very familiar with their work (they're doing a great job), but the demo video you linked appears to be a photogrammetric-based approach. You can tell because highly-textured surfaces are readily mapped, but low-texture regions remain unmapped, despite high coverage by the camera.
Maybe they use learned features for things like persistent AR, but I'm quite certain that they do not use deep learning to predict depth maps ab initio.
I suspect the reason for not using 3D rendering is the desire to cope with the noise and variability of real video.
By using MVS-based approaches, they are able to get over the data hurdle by compiling a dataset of your average YouTube video, instead of creating 3D renderings that include dynamic people. Importantly, MVS is really quite accurate, and in many cases can be considered ground truth.
Being able to forgo 3D renderings to use video only is almost certainly a reason why their results are so good.
Maybe it's the case that this system doesn't actually return an X,Y,Z camera pose, but rather just a pixel specific depth, and not a recovered pose for new inputs.
https://deepmind.com/blog/neural-scene-representation-and-re...
To be fair, this particular application doesn't really need more to show it's improvement over other approaches, but still.
Compared to "Chen et al" which is a bit flickery in the foreground, but full of stable background details, their result is almost completely black 3m in.
I'd still prefer to use explicitly open datasets because it allows for simpler data sharing and easier reproducibility, however in cases where that's not possible whatever is available will do even if I'm restricted in how I can redistribute that data.
They train a DepthCNN to infer depth from monocular images (lidar or stereo for supervision) and make sure it's temporally consistent by adjusting with pixel transformations from the previous and next frame using a PoseCNN. https://arxiv.org/abs/1704.07813
The guys at Google use Optical flow (only previous frame) to make sure their model trained on static object video sequences works when the scene is dynamic using a mask for a specific class an object (humans here). They do have to make sure nothing but humans are dynamic in the scene.