Nice blog post about this stuff: http://willdrevo.com/fingerprinting-and-audio-recognition-wi... - https://github.com/worldveil/dejavu
Nice blog post about this stuff: http://willdrevo.com/fingerprinting-and-audio-recognition-wi... - https://github.com/worldveil/dejavu
http://jack.minardi.org/software/computational-synesthesia/
You can also see the code behind it here:
https://github.com/jminardi/audio_fingerprinting
I am by no means an expert in this area and a few people have since told me I did a few stupid things in my analysis. But you might find it interesting.
Another method to detect an audio pattern is cross correlation on the raw audio signal. But it is very expensive in computation power and memory.
The longest operation with fingerprinting is often the DB query that is associated. Lots of work to do there. In that space, Will Drevo's work is really good. I will share my DB implementation later.
I've always wondered: Is there a way to compare fingerprints with humming sounds or live recordings?
Those fingerprinting techniques don't seem to be suitable for those tasks, do you know of any methods to accomplish this?
If you want to do some research, here is a short review paper on the topic http://www.cs.toronto.edu/~dross/ChandrasekharSharifiRoss_IS...
As for 2d array spectrogram, it is not needed in my lib (expect when plotting is activated). I only care about maxima in the spectrum of each data window. In other words, 1d spectra are enough.
I was mainly responding to the OP's distinction between analyzing a visual representation and analyzing a "2d array" when they are basically the same thing.
This is what I mean. I guess their tooling just outputs graphics and it's easier to work with those than the pure 2d array in numpy or something similar.