Implementing Shazam with Java in a weekend
redcode.nl
redcode.nl
http://yro.slashdot.org/story/10/07/08/2311225/Open-Source-M...
Basically, Shazam sent the author a cease and desist. The post was removed for a bit, but it's now back up.
I did a master degree on pattern matching on text in 1999 (same year Shazam started), and it was obvious then that much of the same concepts (many of them from textbooks stemming from the 1970s) could be used for pattern matching on music. (We even discussed implementing music matching, but couldn't see a market for it.)
There may very well be other parts of music fingerprint algorithms that are patentable, but I have a hard time believing that the parts described in the article could be.
Basically, to get good features you need to find a set of candidate features (points in the image that stand out) and then filter them according to some desirable properties (in standard image recognition, for example, you want them to be resistant to rotation, translation, stretching, etc). There exist very good algorithms for finding these features, which is what Shazam uses. Of course, after finding the features, you need to hash them in a way that is able to match any time in the song, resist noise, etc etc, so it's a very interesting application.
I was very surprised to learn that the best way to recognise a song was to convert it into a picture and then try to recognise that picture. It seems like a roundabout way to do it, but it works very well.
You can find more information in this paper: Viola and Jones, "Rapid object detection using boosted cascade of simple features"
I don't know the details of what they're doing, but most likely they are using something related to SIFT: http://en.wikipedia.org/wiki/Scale-invariant_feature_transfo...
This is another seminal work in computer vision, which solves both problems that 'StavrosK mentioned:
1. Find candidate feature points that stand out, and are reliably and repeatedly
detectable despite image variations.
2. Get a "hash" of each point that can be used to do searches fairly quickly.
Lots of work in detecting and recognizing objects now uses some variant of SIFT, and it's finding usage in lots of other areas of vision as well. I wouldn't be surprised if as many as 10% of papers at the top vision conferences use techniques based on some variant of SIFT.Sometimes, it's way to easy to overlook the obvious.
He also seems to be suffering from 'death by HN', the server is more 'down' than 'up', keep retrying and it will come up.
That Aphex Twin face is on a log scale, not a lin scale, on a log scale it looks much better.
I would imagine the algorithm is different, since the Shazam algorithm is designed to find exact matches corrupted by noise, EQ, etc., and two hummed versions of the same melody may vary in key, tempo, and timbre, and have small rhythmic differences.
It's also nice to see the thinking process he went through when developing his solution to this.
Yes the algorithm might not be perfect, it isn't a Shazam clone, but it does demonstrate that within 48 hours he created something that could recognise music. Now that's a good effort in my eyes!
http://www.popmodernism.org/scrambledhackz/index.html
Has anyone heard anything on this?
;-)