Disclaimer: I work for Google (but not YouTube). Opinions are my own.
I suspect this has to do with economics. YouTube has way too many videos so everything is probably largely automated. Having to go through videos and find all possible modifications to pirated content is really difficult, and it'd be very expensive for them to do both in terms of engineering effort and also computing resources.
After all, it's not like you can search videos by their content, you can only search their titles or descriptions with text. Sure, they could start detecting things that have mirrored content, but then people would just change the method by which they modify the original video. People who are trying to upload pirated content only have to modify the videos they upload, while Google has to check all videos that are uploaded, the majority which are likely not pirated, and checking the videos is likely much more computationally expensive than modifying them a tiny bit.
I admit I know nothing at all about video formats and that stuff so if I'm wrong please correct me as I'm always grateful to be able to learn more.
Err... yes you can. It's trivial.
I would then use these features to identify videos that have the same content as these already reported - pick your favorite clustering algorithm. This would ensure that for a single copyrighted video being reported all plausible manipulations that fall under the model are also discovered.
Point being, this is something that could be implemented in less than a month or two by an individual or small team. If Google hasn't implemented it it's because they don't want to, not because it can't be done.
How do you know this would perform any better than what they are currently doing though? Do you think this would catch the cases which currently pass their filtering like videos with mirrored images? And would this produce less false positives?
According to this post, it seems they already do something similar to what you've described: https://www.quora.com/Does-Content-ID-look-for-a-match-of-th...
I suppose in the end this is all going to just be speculation since they haven't disclosed what they actually do to try to find copyrighted content. Can you even predict how well a machine learning or classification algorithm will do without just trying it empirically?
Mirroring etc. would just be one example of the transforms an approach like this would be robust to.
It's hard to say what the error rates will be, but you should be on the last 5-10% Pd/Pfa pretty much instantly and then if you want more you can just acquire more true labelled data (if a user can find a copyrighted video on your site, I would think a contractor could, too!).
Oftentimes uploaders have to heavily crop or otherwise distort the video to get it to upload without getting caught by the detection system.
I figured they were already checking the videos somehow, but it seems to not catch videos that are mirror imaged or have other modifications. What I was trying to say was that it's much easier and cheaper for videos to be modified than it is to add additional filters through which uploaded content must be screened for copyrighted material.