>
federated learningOne approach that occurred to me is to take advantage of the fact that same ads are played in different TV shows.
Suppose you record hundreds or thousands of hours of TV video/audio, then compare it to your program guide data.
In theory, you should be able to detect that some segments of the video/audio correlate very strongly with what show you're watching (occurring only there), whereas other segments of video/audio appear all over the place in different programs.
And use fingerprinting to avoid having to store/transmit all that video/audio data.
This is not that different from what Shazam does with music recognition, although their job is easier because their audio data is naturally segmented into songs, whereas for ad detection you have to find segments. But I bet there are algorithms to efficiently find common sub-sequences (maybe from DNA bioinformatics?).
Once you have fingerprints of commercials, you distribute a database of these to viewers, and their device uses it to detect commercials while watching.
---
For programs that air more than once, you could invert this and look for segments of video/audio that are common between the airings. The rest is more likely to be commercials, especially if you've got a movie that airs on different networks.
Or if you can see a show/movie through on-demand, maybe you'll get different commercials on repeated viewings.
---
Another approach is facial recognition. If you see Flo from Progressive, it's probably an ad. You can maybe even recognize new, unknown ads this way.
If you want to go really crazy, mine IMDB for TV/movie titles, the cast that appear in each, and their photos, then recognize non-ad video footage based on who you know is supposed to be in that episode.