Remove Moving Objects from Video
medium.com
medium.com
Background:
There are streamers on Twitch who will play copyrighted music as part of their streams. Twitch allows them to do this, but then Twitch will scan these videos for copyrighted music against some music fingerprint database and mute that section of video entirely (the part of the video that played that audio). The other parts are unaffected unless they played some other copyrighted audio. YouTube does something in terms of recognizing copyrighted music, but will demonetize the entire video as a result. Needless to say, demonetization does hurt the streamer.
Alternatives:
I don't want to get into whether or not Twitch and YouTube are right in doing the copyrighted audio matching and the subsequent actions they take. Some streamers who've been affected by this have started playing royalty-free/copyright-free music or music from lesser known artists that are less likely to be in these music fingerprint databases.
My question:
Is it possible to just subtract the audio of a copyrighted track from a video after it has been detected to having being played in a video?
https://towardsdatascience.com/audio-ai-isolating-vocals-fro...
It would be a game changer if someone were to come up with a novel method of decomposing audio into discrete components (e.g. people speaking, specific instruments, background noise).
Likely this would require completely new hardware to capture different audio attributes in addition to simply capturing a stream of vibrations from a microphone.
Edit. If you know the song it should be something simple like do cross correlation of audio with known song. Find peak. Solve for the gain and subtract away scaled and shifted song from original track. Will be rubbish if gain and timing have errors. Might need to do it in little chunks and interpolate the gain and shifts.
Edit 2. More generally, you might want to worry about the song having passed through some unknown transfer function (i.e. it is being played and recorded through shitty equipment). Then you have an interesting inverse problem. If everything is linear it will involve a regularized deconvolution. Will be tricky then.
> It would be a game changer if someone were to come up with a novel method of decomposing audio into discrete components
It's something that has been generally addressed and ~works. It will obviously depend on the specifics of the application, and yes if you can constrain the problem space further you ought to do better!
Someone mentioned the software that can isolate instruments from one another. But the problem is this usually relies upon the instruments or vocals having a different position in a stereo signal, and different frequency ranges, to perform that extraction. This is much easer to do in such cases, especially since the material you are working with is generally clean source material. Here you have a mixed signal, that may or may not contain stereo data, but certainly contains a lot of background noise, as well as other game audio you wish to retain.
What is needed is a far more intelligent method of identifying the spectral signature of a known source by matching the target with the source in the time domain, and then using that identified signal in the target itself to cancel itself out.
In practice this would mean comparing the amplitude modulation of each spectral band in the source to the target to identify the spectral components that need to be removed, and then using the corresponding "matched" bands, generate a signal, using the source as a guide, from the target audio itself that represents only that portion of the spectral signature that is being matched to the source. This newly generated signal can then be used to remove that audio from the target with 180º phase cancellation.
Keep in mind that wherever the spectral bands overlap with other audio content you wish to retain, there will be a lot of artifacting and signal loss or phasing. A second pass would need to use the equivalent of inpainting to reconstruct those missing components.
Overall, it's a very hard problem.
Of course those are not open-source but it's really inspiring to see such uses of AI. Another very interesting one when it comes to video has been [2]
[1] https://www.youtube.com/watch?v=25ltIoHtiO4 [2] https://github.com/avinashpaliwal/Super-SloMo
I've used such a filter to remove dirt and dust from 8mm films via avisynth, a program that goes back more than 20 years, though the filter in question is not quite that old -- at least 10 years.
Here is the filter in question: http://avisynth.nl/index.php/RemoveDirt
It still looks amazing of course