Shazam was founded in 1999-2000 ish and was live by 2002, longer ago than many people realise. The fingerprinting method they used was a novel one, and they successfully deployed it in a couple of years, with a database of two million tracks, in a system that could be used by dialling a number from any phone. It's an extraordinary and rare example of taking a research method and getting it to work at scale in the face of real constraints, in the process producing something that most people would never have imagined could be done.
Would really be interested if someone has tried other approaches and can comment.
This is a really nice paper as well, but the Shazam method is more exciting because of its focus on the search part of the problem. Audio fingerprinting wasn't an entirely new field of course - the Philips paper also cites several prior publications.
The big limitation of all these methods is that they are robust to EQ and other degradations but not at all to variations in timing between frames. So they work well for identifying different instances of the same recording, but not at all for matching different performances, or for "query-by-humming".
https://github.com/jvbalen/soundofshazam/blob/master/sound_o...
And an example output (try Shazam-ing it!):
He also got some legal threats just for publishing code he created himself.
[0] http://royvanrijn.com/blog/2010/06/creating-shazam-in-java/
The remaining thing I don't understand now is how the heck they managed to negotiate the license for all the songs they have in the database. A friend of mine had a major website for lyrics, and he got some legal trouble from the record companies. When he tried to buy a license to display the lyrics, after some back and forth, they basically said they're not interested in licensing. This was in the early 2000s before streaming took off. Since that time I have the impression that often legal and social problems are much harder to solve than technological ones.
I actually bought Shazam on android after repeated attempts at getting google to find music in the same way - they're laughably bad.
that said, i still prefer shazam because it's an actual app, whereas google's implementation is a button that sometimes appears in the google app, and sometimes doesn't, seemingly at random.
Not really "magic" how it works but still pretty amazing that I can pull the phone out of my pocket any time I hear a song I like and immediately see the band/title.
Is it on by default? Cos that sounds really creepy.
[1] https://motherboard.vice.com/en_us/article/8q8ee3/shazam-kee...
That's really sweet—I wonder how much storage that requires? I would have expected it to require a lot, but if it's on a phone, then maybe not that much?
For anyone else finding this as implausible as I was, apparently the Pixel has a local database of 17,300 tracks. Still pretty impressive though.
https://uk.pcmag.com/news/91606/googles-pixel-2-phones-recog...
https://ai.googleblog.com/2018/09/googles-next-generation-mu...
It's also pretty garbage for classical music.
(o) https://en.wikipedia.org/wiki/Shazam_%28application%29#Early...
I have worked on integrating a shazam like library inside an app. We have looks at more than 12 solutions before finding a company with a song recognition library good enough to be called shazam like. It is by no mean a small feat you can reproduce during a weekend.
The interesting thing is: most of these songs are on YouTube, so in theory Google could build a better Shazam.
I think it's a good thing, because often there are good comments in response to downvoted comments. Letting people delete comments would make the discussion hard to follow.
It was informative to myself at least.
Separately: there was nothing wrong with gnulinux's comment. It was obviously posted for the pleasure of sharing information, and didn't put down either the GP or Shazam.
I wish someone would do the same thing for commercials on TV. Auto detect the commercial and advance the DVR by its exact duration
They've had it for years now. I was skipping commercials on my HTPC back in 2007-ish. I think the older Tivo's also had this ability.
AFAIK that's not the case of Shazam where a read-only database would be enough. Sure eventually you have to manage multiple versions of your app and be able to gradually update your DB, but that's still simple stuff compared to the problem of how to make money.
I've never used Shazam, but I used a similar feature on my LG Chocolate back in the day. How is Shazam any different than what my LG Chocolate did (aside from the Spotify integration)?
Expecting people to know how is it different from an LG Chocolate is just absurd...
That said, Shazam is actually older, having started as a shortcode service in the UK way back in 2002.
How exactly did this come to be?
That's the anthology I want to read.
I remember Adam Neely (A musician/jazz composer/youtuber) mentioning in a Q&A, that for songs they played live for a television broadcast, backing tracks were added underneath and it was mixed as close as possible to the originals, among other reasons for the purpose of them being better "shazamable".
What are you referring to here?
It doesn't recognize you singing because it doesn't have the same spectral content as the full song. Shaman doesn't really identify tunes, it identifies recordings. (And it is amazing to me that it works at all)