Audio Fingerprinting with Python and Numpy (2013)
willdrevo.com
willdrevo.com
https://linux.ucla.edu/~eqy/molasses.html
The biggest surprise for me was that the entire process took only a few days on a single laptop (Sandy Bridge, no SSD).
That's almost a year of round the clock music playing 24/7 for free and completely unencumbered by any sort of licensing or playback restrictions. Most modern media players will play the formats back also without trouble (VLC for example).
Awesome write up, I'm looking forward to reading it all.
Huge thanks to Deja-Vu! Awesome library
I get a couple emails a week about it, but probably the weirdest/coolest was a guy in Spain who used Dejavu to make two rap-dueling robots who spoke in Basque.
Actually, I encourage anyone interested to look up Panako[1], which is meant to be a framework for comparison of different audio fingerprinting techniques.
it loads the page fine, but two seconds after that it crashes
https://en.wikipedia.org/wiki/Shazam_(service)#Patent_infrin...
The core algorithm is well known and fun to implement:
https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf
A blog had a simple implementation of it and were contacted by Shazam's lawyers:
Pretty sure Shazam/Philips/Dolby patents cover Europe.
Contrast this to now with you would like to know what artist performs a song and you Shazam a song in one app and then when you identify it you open up Deezer or Spotify or whatever and then have to type it in the search box and then add it to your library or a playlist.
By integrating fingerprinting you can identify and add new music all in the same app.
You only truly understand something when you can explain it in lay (or close to) terms, I find.
The American vernacular is you get "dejavu" whenever you have that "I've been here before..." feeling.
Dejavu uses a locally sensitive hash (LSH) just like any other approximate search hash might. The key to note is that we're binning both the frequency and time units of the spectrogram, giving us room for noise/error. In fact, you can tune the granularity to which this happens by adjusting the Dejavu FFT window (DEFAULT_WINDOW_SIZE). This will create a spectrogram with few (and therefore larger) cells.
The trade-off with smaller cells is that with too much granularity (or if the audio is even a tiny bit stretched or we have small Doplar shift effects), we may miss the fingerprints we want (false negative). On the other hand, with too low granularity, we risk having our frequency bins too large and having perhaps another song/query audio match when it shouldn't (false positive).
Luckily in either case, we don't need to see all the fingerprints, just enough to align properly in time.
So to answer your question, Dejavu is fairly resistant to noise (and can be tuned with FFT settings) as long as the audio's original timing is unchanged.
Most people don't realize this, but usually artists/labels will have a different mastered version of a song on each platform. Spotify, for instance, has it's own normalization algorithm that it puts tracks through to even out the listening experience in terms of loudness (RMS). Of course artists and their mastering engineers want to have some control over how that sounds, so they will change it.
One major problem with video is that you could have SD/HD versions and/or full frame/original aspect ratio types of differences of the same movie. One idea I wanted to play with was to detect edit points. The number of frames between edits could be used as the fingerprints. The entire concept was never anything more than a thought exercise. For the purpose of the exercise, we had to assume that there is no audio with the picture.
There are a lot of FFT libraries to process image data since most image compression techniques use some sort of FFT. Would this same type of fingerprinting be able to be used for a visual image. Could the amplitude of the RGB frequencies be used over time? The data set would increase with 3 channels of color, but would it also not help decrease false positives by making the combinations more unique?
[1]: http://www.hackerfactor.com/blog/?/archives/432-Looks-Like-I...
The EP fingerprints were a lot smaller, though. IIRC it used a mix of beat and tonal detection.
The other interesting differences to note are that Echoprint doesn't use a constellation fingerprinting approach along with offsets, and the fingerprinting is meant to be the same across all platforms / use cases so you can compare them.
As a direct result, you also can't get the offset in seconds that your query audio refers to like you can with Dejavu.
When I coded up this project, I wanted something that was more customizable - allowing you to decide the speed, number of fingerprints, size of the fingerprints etc to match your own false positive / memory / CPU requirements.
When you do, you sacrifice interoperability between all Dejavu index installations, but you gain that application specific performance. It of course depends on your use-case which library is better.
It didn't end up becoming a telco service solely due to commercial agreements, but it was a lot of fun and almost embarrassingly accurate with ABBA songs (since we ended up trying a lot of variations on the first entries in the catalogue).
You could, for instance, insert other tracks into the audio additively to try and confuse the fingerprint retrieval logic into suspecting a different track, but since this and many other fingerprinting techniques depend on the actual frequency of the audio emitted, there's no shortcut to obscuring the actual track.