How Shazam Works (2015)
coding-geek.com
coding-geek.com
I used it back in the day and I recall it working very well.
There's also a paper on the technology [2]
[1] https://github.com/search?utf8=%E2%9C%93&q=echoprint&type=
[2] http://mediatechnology.leiden.edu/images/uploads/docs/wt2015...
Shazam has always been a pretty amazing product, especially when you consider that it was first released nearly 20 years ago. It has also been fairly poorly understood, so I am glad that people are taking the time to revisit and review the technology behind it.
The cepstrum is good for providing energy compact/sparse representation of harmonic features. This is why it's used (/was used) a lot in speech recognition. Speech sounds tend to have harmonic properties (see: formants). Frequency-based transforms tell you how much of a frequency (repeating/periodic signal) is present. If you have harmonics, those can sooort of be thought of has repeating patterns in the frequency domain. So taking a frequency transform of a frequency transform (which is super loosely a cepstrum) gets you a nice compact (separable) representation of inputs that tend to have harmonic features.
Most music tends to be pretty damn harmonic... so maybe?
Also an argument to be made that if you have a big enough network of the right kind that's just taking in windowed time domain data (almost surely involving recurrence), it might not be surprising that you could find some cepstral-like stuff naturally pop out.