Deblur-GS: 3D Gaussian splatting from camera motion blurred images
chaphlagical.icu
chaphlagical.icu
However, I’m not sure how admissible this would be as evidence given the somewhat generative nature. I would assume any lawyer would tear apart the fact that this is guesswork/estimation and not a known “truth”. Deblurring tech already exists to mathematically undo motion blur, but gaussian splatting would effectively be “creating” evidence.
Allowing generated images like that to be considered evidence of crimes would be an incredibly dangerous precedent to set, in my opinion. But with this court, and the right case, dangerous, unpopular, and otherwise questionable precedents seem to be the name of the game these days. Especially when judges can be bought off without any repercussions.
Down for me, archive above.
It seems to me they don't use any ML at all. They use backpropagation to jointly optimise the entire physics/motion model, which models camera motion and the generated blurry images (they generate multiple images for each camera frame along the path of motion of the camera, and then merge them, simulating motion blur)
The data they optimise over is just the images of the current camera trajectory (as far as I understand)
Here's some early work in this area which seems promising: https://guanjunwu.github.io/4dgs/
The OP paper is cool but isn't alone, here's some concurrent work: https://github.com/SpectacularAI/3dgs-deblur
Also related from a couple years ago, using NeRF methods (another area of current 3D research) to denoise night images and recover HDR: https://bmild.github.io/rawnerf/ NeRF, like Gaussian Splatting, seeks to reconstruct the scene in 3D, and RawNeRF adapts the approach to deal with noisy images as well as large exposure variation.
In terms of Gaussian Splats vs GenAI, usually GenAI models have been trained on a prior of millions of images so that they can impute / inference some part of the 3D scene or some part of the input images. However Gaussian Splats (and NeRF) lack those priors.
Video, where the result needs to be temporally coherent and make sense in 3D, can't be the easier one.
Why not? Video is a much more tractable problem because you have much more information to go on.
When you want to stay faithful to the actual data then your options are limited, for quite a large part of the image a simple convolution is about as good as it gets, except for the edges. Basically the only problem we couldn't solve 20 years ago was excessive ringing (which is why softer scaling algorithms were preferred). You can put quite a lot of effort into getting clearer edges, especially for thin lines, but for most content you can't expect too much more sharpness than what the basic methods get you.
And then there is the generative approach where you just make stuff up. It's quite effective but a bit iffy. It's fine for entertainment but it's debatable if the result is actually a true rescale of the image (and if you average over the distribution of possible images the result is too soft again).
In theory video can do better by merging several frames of the same content.
Note that a limitation of this result is that it assumes a static scene, but that's already a typical limitation of most gaussian splat applications anyway, so it kind of doesn't matter?
I'm trying to make an opensource NN camera pipeline (objective is to be able to run on smartphones but offline, not real-time), and I'm still barely managing the demosaicing part... would you be open to discussing with me?