I don’t think this is a novel idea, but it is still a great topic for a PhD. While the results in this paper look impressive, my suspicion is that the system doesn’t generalize particularly well. (I suspect this from experience with similar, albeit simpler, ideas, as well as from looking at the datasets.) If you can make a system that generalizes to new environments and objects, or a system that works with real-world natural image/video data, that would be a tremendous accomplishment.