Website: http://sniklaus.com/kenburns
Website: http://sniklaus.com/kenburns
So 500 GB are about 2800 full video downloads.
/kenburns/promo.mp4 | GET | 175574 | 33.4183748571803
That the video was watched 175574 times? That would indeed be a lot!
Hope you're able to release it ASAP!
I skimmed the paper about the Ken Burns effect, and if you don't mind I have some questions. I hope I didn't miss the answers in the paper itself, I'll be sure to read it more carefully when time permits.
1. The loss function for the depth estimation is Ldepth=0.0001·Lord+Lgrad. Is Lgrad much bigger than Lord by design or will this basically make Lord tiny and almost unecessary? How did you arrive at this number and not a magnitude bigger or smaller?
2. How are you rendering the point cloud, like tiny discs in free space? When I think of point clouds I think of the typical LIDAR output renderings, but your results are continuous images, or video frames. Are the points rendered to an image plane and the result then interpolated?
And a couple more general ones:
1. Are all other depth estimation techniques other than NN obsolete now? Is there no point in estimating intrinsic camera parameters and epipolar lines when the state of the art seems to be to input one or more images into a NN and let it produce the depth for you?
2. How do you decide how the NN should look like, with various downsampling layers, convolutions, etc. I've seen people start with pretrained networks and retrain them, but how do you build a novel network based on the input and desired output?
1. I am not sure about the scale of the individual loss functions anymore, my apologies. I determined the combination of the two losses via a simple grid search and seeing what works best (plus / minus a magnitude did not make that much of a difference).
2. The points are just splatted to an image plane, more advanced point cloud rendering techniques would be better though. There is no video frame interpolation, each individual frame in the output video is a rendering of the point cloud from a different camera perspective.
3. I am not sure about multi-view stereo, COLMAP still seems like the state of the art for that. But neural networks definitely outperform classic techniques for single image depth estimation.
4. Common architectures just did not do as well as I was hoping for so I tried about 1500 model architectures. I started with an architecture that intuitively seemed right and then gradually explored / refined alterations of it. It ultimately was a lot of trial and error.
Was it the university that didn't want to release it? Are they looking at commercializing it, or how does that work? Is it available in any commercial software? It kind of looks like magic and would probably be very useful for a lot of purposes.
2. So basically each point is projected to the image plane without perspective mapping? So in 3D, the further away from the camera they are the bigger they are so they all have the same size on the image? And that prevents any seams to occur in the pixel grid as things move around?
4. Experience, intuition, and elbow grease. Kind of what I thought, but I guess it's reassuring to see an expert in the field having to try 1500 variants.
2. Yes and there are two mechanism for handling seams. First, the inpainting which extends the point cloud and can provide a higher sample rate. Second, a postprocessing step that heuristically fills in any seams that may still be present despite inpainting.
4. The downside of it is that one needs a lot of resources in order to try all of these variants, which not everyone is lucky enough to have access to.