the original paper is quite readable: https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf
shazam's value is obviously in how it scaled this method to millions of users and songs but implementing it for yourself on a limited catalog of songs is a couple days of work once you have the theory. in fact this was a lab project for the intro signal processing class at berkeley that i ta'd.