Texture Enhancement for Video Super-Resolution
github.com
github.com
Upshot: Event Cameras are a different sort of camera in that they have an array of sensor pixels, and sensors only fire when there is a brightness change for that sensor. This has a bunch of benefits, including very high dynamic range, reduced ghosting, and high frame rates, and has some downsides, like reconstructing video, and presumably others.
The paper seems to have started out with the idea that if you had event camera output, you’d be able to reconstruct more fine texture details. And, this works incredibly well, their baby model trained for 8 days significantly beats SOTA and looks a lot better in comparisons as well.
They then seem to have added a step where you simulate/infer event camera data from “normal” RGB video, using a different set of networks, and use that inferred event data to do the texture recovery, and … this also works.
Pretty surprising, and interesting. Their GitHub is full of people like “I want to try this” and then realizing it’s a fairly deep stack to deploy. Even as is, it seems worth someone building a GUI around this in an app, it’s quite remarkable.
If someone else manages to deploy and try this please share your result.
That said I don’t understand it very well, for instance there’s a voxel step in the pipeline and I have no idea why.
Upon close inspection the plate's digits look realistic, but there are some symbols that look unfamiliar to me. But I don't know what country the footage is from, so I don't know real from unreal when it comes to symbols.
If the car's badge turned out to correctly match the car model, that might be a bit of a red flag. Although it's not out of the question that a model could eventually recognize car models and get badges right. It just seems unlikely that I'd see such an advancement in a video before I ever saw it in a still image.
It is the case that with completely generative models, you will get hallucinated details very likely to be untruthful. But with this approach, you can see blurry input images of license plates that with our naked eye we could not possibly decipher the characters, then put through this model where the output is very close to the actual ground truth.
https://dachunkai.github.io/evtexture.github.io/static/image...
Where does this information come from? It seems they are generating synthetic "event camera-like" events just from diffs between still frames? So maybe they trained a model based on real events from a real event camera? It's hard to tell from their write-up. But these results are very impressive.
Then there is undoing reversible transforms, such as some blurs. That makes information that was there all along more legible. Such as the example you have there.
This paper is a case of both. It does upscaling, but it uses temporal information to find additional constraints that can be used to restrict the degrees of freedom of the "making values up" part. So it's part information recovery, part hallucination.
https://github.com/uzh-rpg/rpg_vid2e?tab=readme-ov-file#read...
The video data is simply a stream of events which encode the time and location of a brightness change. For an immediate full-scene change (like removing the lens cap), you’d get a stream that happens to update every pixel, but there’s no particular guarantee about the ordering.