ReconFusion: 3D Reconstruction with Diffusion Priors
reconfusion.github.io
reconfusion.github.io
In effect, I believe what happens here is that where the actual data is ambiguous, they use the diffusion model's hallucination ability to fill in details (which, thus, weren't in the training data).
But since diffusion models are known to memorise things (like the Getty watermark) and were trained on most of the internet, there is also a chance that the diffusion model memorised an actual frame of the original training data which was withheld during evaluation. So there is a chance of groundtruth data leaking into the testing results, I believe.
From their paper:
"Our base diffusion model is a re-implementation of the Latent Diffusion Model [42] that has been trained on an internal dataset of image-text pairs with input resolution 512×512×3 "
They also use multi-view datasets during training, but presumably they haven't included those in the diffusion pretraining.
"Training Dataset To learn a generalizable diffusion prior for novel view synthesis, we train on a mixture of the synthetic Objaverse [10] dataset and three real-world datasets: CO3D [38], MVImgNet [64], and RealEstate10K"
to me sounds like the diffusion model had access to 87% of all frames for it to memorize. Then reconstructing from 3 views + diffusion is closer to using those 3 views to recall the memorized 100 views and use those for reconstruction.