Real-Time Coherent 3D Reconstruction from Monocular Video
zju3dv.github.io
zju3dv.github.io
This is the only video I could find, but we were doing monocular reconstruction from a limited number of RGB (not depth) images AND doing voxel segmentation on the processing side. https://www.youtube.com/watch?v=nqy44VSWh3g
Even as far back as 2010 people were doing reasonable monocular reconstruction including software like meshroom etc...the whole of TU Munich also under Matthias Niessner has been doing this for a while.
What's novel here?
At a minimum though 6D.ai and a few others had companies that were selling this as a service at least as far back as 2017.
I don't recall you having similar real-time meshing functionality in 2016-2017, Andrew. Can you show what you had?
As far as I'm aware, Abound was the first to demo real-time monocular mobile meshing: on Android in early 2017 (e.g. https://www.youtube.com/watch?v=K9CpT-sy7HE), and iOS in early 2018 (e.g. https://twitter.com/nobbis/status/972298968574013440).
The actual 3D reconstruction is so-so, I agree. And they kinda cheat by using ARKit (which uses LIDAR internally) to get good camera poses even if there is little texture.
So the novel part here is that they can immediately merge all the images into a coherent representation of the 3D space, as opposed to first doing bundle adjustment, then doing pairwise depth matching, then doing streak-based depth matching, and then merging the resulting point clouds.
Also, they can use learned 3D shape priors to improve their results. Basically that means "if there is no visible gap, assume the surface is flat". But AFAIK, that's not new.
EDIT: My main criticism of this paper after looking at the source code a bit would be that due to the TSDF, which is like a 3D voxel grid, they need insane amounts of GPU memory, or else the scenes either need to be very small or low resolution. That is most likely also the reason why they reconstruction looks so cartoon-like and is smooth on all corners: They lack the memory to store more high-frequency details.
EDIT2: Mainly, it looks like they managed to reduce the GPU memory consumption of Atlas [1] which is why they can reconstruct larger areas and/or higher resolution. But it's still far less detail than Colmap [2].
TSDF memory isn’t an issue since Niessner et al. (2013).
It doesn't by default, for power reasons, but it will in a pinch.
Would love to know in which circumstances it’s used. I assume you work for Apple to know this so understand if you can’t share more.
I would strongly disagree. This paper uses TSDF and runs into memory issues. And ATLAS is using TSDF and running into memory issues. So for practical applications, TSDF is still too memory hungry.
Try out our app, Metascan, to see an example of using TSDF with a multi-resolution GPU hashtable that only stores voxel data near surfaces. Or just skim the original voxel hashing paper from 2013 to understand the technique.
Storing voxel data in an array is a lot simpler. So if it’s not the focus of the research, then why would academics engineer something more complex?
Of course, this network may do well on this because it's trained on indoor scenes with walls of uniform height and width. Uniform implies flat is likely to emerge as an assumption. You get to see the wall/floor joints and the wall/wall joints, so you have some references.
Remember the Tesla that hit the big white semitrailer because the algorithm couldn't measure depth to a uniform surface? This is a hard problem in unstructured situations.
First one was this but it was slow: https://www.matthewtancik.com/nerf
Then it's got faster: https://www.youtube.com/watch?v=fvXOjV7EHbk
Lot of interesting papers:
This seems to fit into the genealogy of KinectFusion, ElasticFusion, BundleFusion, etc.
https://www.microsoft.com/en-us/research/wp-content/uploads/... https://www.imperial.ac.uk/dyson-robotics-lab/downloads/elas... https://graphics.stanford.edu/projects/bundlefusion/
Very impressive work. I have not seen any use cases for online 3D reconstruction unfortunately. 6D.ai made terrific progress in this tech but also could not find great use cases for online reconstruction and ended up having to sell to Niantic.
Seems like what people want, if they want 3D reconstruction, is extremely high-fidelity scans (a la Matterport) and are willing to wait for the model. Unfortunately TSDF approach create a "slimey" end look which isn't usually what people are after if they want an accurate 3D reconstruction.
It SEEMS like online 3D reconstruction would be helpful, but I have yet to see a use case for "online"...
Also youtube videos and such with more complex animated characters jumping around and things.
If they could convert that to a floorplan diagram it would be wonderful for UI design, lots of use cases for a map.
San Francisco's recycling operation, at Pier 95, uses a system built by them.
It's not a glamorous technology, but it gets the job done.
Also, to measure the performance / evaluate observations generated from this tech, you would want to compare it to a pretty sizable 3D ground truth set which Tesla does not currently have. There are pretty big advantages to starting with a maximal set of sensors even if (eventually) breakthroughs turn them into unnecessary crutches.
That's not true in the case of Tesla. They started as a EV company that offered autonomy later and since they were selling the product, they had to decide a minimal set of sensors that would, in theory, still work to keep the product cost reasonable.
Tesla does have a sizeable 3D ground truth; they collected tons of data mounting LiDAR on their test vehicles.
"Sizable" for evaluating safety has to be big enough to give small confidence intervals on your error -- and with self-driving, you need to cover a robust set of rare scenarios upsampled from general driving as well. I really doubt Tesla has what they need to convince themselves of higher levels of safety.
Monocular 3D reconstruction can require many frames at many angles (or possibly just a few frames at sparse angles).
The inference case with a self-driving vehicle may allow for this in some scenes, but certainly not all. Trying to infer the relative motion of say two moving vehicles and getting enough frames for monocular reconstruction may take longer than is required for an emergency break, not to mention the robustness issues with those pose estimations without additional sensors. Solving any one of these issues can be done but I think it’s pretty clear we’re a bit away from solving all of them.