Show HN: Shazam-like acoustic fingerprinting of continuous audio streams
github.com
github.com
This lib is a brick in an adblock for radio broadcasts I have been developing for a while and that I am progressively open sourcing.
Edit: I don't want to support the notion that we should avoid 10$/month for such a great service, I was just curious about the technical implementation.
For a free Spotify without ads, have a look at http://www.stationripper.com/ (it's old software)
Edit: Station Ripper had caused a law/politics debate in 2005 in France about the right to do private copies of legal media http://www.assemblee-nationale.fr/12/amendements/1206/120600...
I figure any block of audio longer than a few seconds that appears in more than one podcast episode could be snipped.
I think it would have to be desktop software since if it were a service and edited the podcast files then it would be copyright infringement. I guess it could be a podcast player and not actually distribute the edited files but just the timestamps to skip ads.
Personally I'd like software I can run over a directory of mp3s that would remove duplicate sections. Any thoughts on feasibility? I'm surprised it hasn't been done yet.
Edit: https://github.com/ppwwyyxx/speaker-recognition/ looks like the first of a number of good starting points.
My understanding is that Apple & Spotify have signed up to these guys with a view to correct payments for artists with user uploaded mixes [1].
[1] http://variety.com/2016/digital/news/spotify-apple-music-rem...
https://venturebeat.com/2017/10/19/how-googles-pixel-2-now-p...
Thanks a lot
discussion: https://news.ycombinator.com/item?id=15619416
supported devices: https://wiki.lineageos.org/devices
Personally I've only ever seen "Pixel Ambient Services" show up on my battery list once, and that was 1% usage after a fairly long day out.
Nice blog post about this stuff: http://willdrevo.com/fingerprinting-and-audio-recognition-wi... - https://github.com/worldveil/dejavu
Another method to detect an audio pattern is cross correlation on the raw audio signal. But it is very expensive in computation power and memory.
The longest operation with fingerprinting is often the DB query that is associated. Lots of work to do there. In that space, Will Drevo's work is really good. I will share my DB implementation later.
I've always wondered: Is there a way to compare fingerprints with humming sounds or live recordings?
Those fingerprinting techniques don't seem to be suitable for those tasks, do you know of any methods to accomplish this?
If you want to do some research, here is a short review paper on the topic http://www.cs.toronto.edu/~dross/ChandrasekharSharifiRoss_IS...
As for 2d array spectrogram, it is not needed in my lib (expect when plotting is activated). I only care about maxima in the spectrum of each data window. In other words, 1d spectra are enough.
I was mainly responding to the OP's distinction between analyzing a visual representation and analyzing a "2d array" when they are basically the same thing.
This is what I mean. I guess their tooling just outputs graphics and it's easier to work with those than the pure 2d array in numpy or something similar.
http://jack.minardi.org/software/computational-synesthesia/
You can also see the code behind it here:
https://github.com/jminardi/audio_fingerprinting
I am by no means an expert in this area and a few people have since told me I did a few stupid things in my analysis. But you might find it interesting.
[1] https://github.com/JorenSix/Panako
[2] http://www.terasoft.com.tw/conf/ismir2014/proceedings/T048_1...
https://www.youtube.com/watch?v=K6FxfZH_ZK4
The phone in that video is just playing a song, it doesn't have any connection to the computer at all.
I'm wondering how to use my Chord Progression data to make a different audio fingerprinting algorithm.
This lib is a brick in an adblock for radio broadcasts I have been developing for a while and that I am progressively open sourcing.
e.g. two independent streams, identifying the same 30-second commercial, but the audio streams are offset from each other by half a sample length?
Maybe some fingerprints will be present in only one of the two streams, but most of them will be present in both.
BTW a few months ago, I talked to an Australian dev that did podblocker.com, but the project does not seem active anymore
Commercial services are available in that field. ACRCloud was mentioned in another comment.
I'm in France and this lib is software only, so probably Shazam patents are not enforcable here.
Anyway, IANAL and cheers to Shazam people
[1] http://www.dubset.com/mixscan/#intro-2 [2] http://variety.com/2016/digital/news/spotify-apple-music-rem...