Self-Supervised Video Object Segmentation by Motion Grouping
charigyang.github.io
charigyang.github.io
https://research.fb.com/wp-content/uploads/2020/07/Passthrou...
The proposed self-supervised approach is comparable to those top methods, even without using RGB, and any manual annotations.
In terms of runtime, it should not matter. Generally speaking though, the overhead of optical flow is often overlooked. For video DL applications, optical flow calculation often takes more time than inference itself. For academic purposes, datasets are often already preprocessed and the optical flow runtime is not mentioned. Doing real time video analysis with optical flow is quite impractical though.
This person asked about clustering and segmenting into more than two separate groups.
Tagger: Deep Unsupervised Perceptual Grouping
What proposed in this motion grouping paper, is more like on the idea level, which gives an observation that, although objects in natural videos or images are of very complicated texture, and there is no reason a network can group these pixels together if no supervision is provided.
However, in motion space, pixels moving together form an homogeneous field, and luckily, from psychology, we know that any parts of the objects tend to move together.
No code as yet, but looking forward to having it released and playing with it: https://github.com/charigyang/motiongrouping
*Edit: It will be interesting to see if this works Video > Image as well. Much of the current video AI work stems from the image foundation, ie split a video into constituent frames and run detection models on those images. But the image detection/segmentation models assume each image is different, and so the process when parsing video this way is unnecessarily complex - sequential video frames in the same scene are more alike than different.
If good segmentation models for video can be more easily trained using this method, then it would be interesting if they can also be applied accurately to still images, since a snapshot of a video is a single image anyway.
The primary challenge I can think of is that the different network structure required makes transfer learning a feature extraction backbone trained from imagenet etc difficult
This looks interesting. Is there any other good works in this area?