Self-Supervised Tracking via Video Colorization
ai.googleblog.com
ai.googleblog.com
The only complaint I have is that it's not better than supervised object tracking, so I wonder if this idea is too late?
To draw a parallel to image classification, at one point in time neural nets were trained with a bunch of unsupervised pre-training using reconstruction loss, but that technique has basically fallen by the wayside as we've gotten larger datasets and found a pile of tricks for training them from scratch.
I'm pretty sure that making it cheap to use new datasets is very valuable in the long run
https://medium.com/gifs-ai/interactive-segmentation-with-con...
http://webee.technion.ac.il/people/anat.levin/papers/coloriz...
The results in their videos look very poor.
They are for some reason saying they can track things with their colorization, when their colorization is extremely unimpressive as well as the tracking that results from using it.
There is no reason colorization needs to happen to do the tracking anyway. The tracking is unimpressive and now indirect.
This isn't some sort of epiphany they've discovered, they are just reinventing video image segmentation very poorly.
Here are half a dozen examples from a 30 second google search:
https://www.youtube.com/watch?v=juDvLrFQF0U
https://www.youtube.com/watch?v=JYgyDdLf7GQ
https://static.googleusercontent.com/media/research.google.c...
https://perso.liris.cnrs.fr/nicolas.bonneel/InteractiveMulti...
http://files.is.tue.mpg.de/black/papers/TsaiCVPR2016.pdf
https://graphics.ethz.ch/~perazzif/bvs/files/bvs.pdf
The only reason this is news is because it's google and the researchers seem to think they've discovered something. Techniques like this with much better results have been shown at Siggraph for decades.
Same is true actually for many DL papers. They'd be actually cool, if they weren't oversold.
As for overselling, I'd say that it's somewhat standard writing style in CS academia to oversell[1] but this is hardly an example of that.
[1] It'd be unusual, but admittedly refreshing for a paper to say "this is just an incremental tweak on existing methods" or something along those lines.
It is as if defending or advertising "deep learning" was the purpose of the paper. It is not. The purpose of a paper is to show a solution to a problem. Much of DL literature (again not all) is a "solution in a desperate search of a problem" rather than the opposite.
I think many of these papers (including this one) would make a great blog post, but just isn't quite enough in terms of scientific content for a full blown paper. A curiosity, nice gimmick, but nothing more. Not really a solution to a problem, not really any idea of non trivial universality.
As the paper itself states, the tracking results are not the absolute state of the art, but they are in the same ballpark, and more importantly, learned without supervision - just watching video. This makes it easier to train on whatever dataset you might have lying around, and more importantly, it's a clever, simple idea that can be improved on and adapted for different tasks.
(Disclaimer: Authors are acquaintances of mine.)
Of course colorization tracks to objects, it wouldn't work if it didn't.
This is essentially automatic video image segmentation, which itself is heavily derived from and related to natural image matting.
Natural image matting could even be seen to be a combination of clustering and somehow solving (or minimizing) the error in the matting equation described by porter duffman compositing algebra.
So, automatic video segmentation can be seen as clustering over 3 dimensions of pixels - x, y and time, with some loose expectations of coherency over time.
There are many ways to achieve this, which should be obvious if you watch some of the videos or glance at some of the papers I've linked.
One simple way is with a bilateral filter to iterate over the volume of pixels, which gradually clusters them together. One of the papers shows this technique.
Everything I linked gives much better results. They don't require 'deep learning' and the idea that colorization follows objects is so trivial it's nonsense to make a paper out of it. This is more a case of visibility and most people not knowing the research that has already been done. That's understandable for people here, but the authors of this paper should have known better.