It fits into that space of an unusual but comprehensible problem, unlike superficially similar features like recognizing animals or objects in images, which is mostly weird ML magic.
Rather, matching two recordings of the exact same performance (one ingested by Shazam at training time and one ingested by Shazam at run time) is more akin to identifying individuals (facial recognition) than identifying species.
You're dead on that it's pretty difficult if you don't benefit from others, we did a ton of work that in retrospect wasn't necessary. I liked the advanced psychoacoustic model, faithfully implemented in high performant C direct from Zwicker. (Psychoacoustics). To a first approximation, about 10/s model -> pca -> top 16 dim -> VQ and the resulting bytes contain more than 50% of the entropy (!!) Shove all of those in a home grown what-you-now-call-a vector DB, do dozens of range queries, and search for any song common to multiple results. Boom, music recognition. Understandable in retrospect but things like that aren't Everest they're like... multiple unclimbed mountains.
0. And far too early to have any applications. Company existed 2000-2001 \o/
https://patents.google.com/patent/US7853664B1/
https://patents.google.com/patent/US6941275
Very interesting hearing about all of the differing approaches people have taken to solving this problem! Do you have further writings on this topic?
I'd argue that Shazam doesn't have ads, rather it is an ad. You search for the song, then see links to buy it in Apple Music. You'll also see "subscribe to Apple Music" type widgets on just about every screen on the app.
What Soundhound does these days that Shazam doesn't (I think; I haven't actually tried Shazam in a long time) is that it displays lyrics for many songs, and is often able to synchronize those lyrics with where you are in the song.
It makes more sense if you think of a production like AGT less as the reality show it pretends to be, and more as a promotional reel for labels.
Of course the content they choose to promote is indexed.
> Alphonso's software uses the same technology that Shazam and similar services employ to automatically detect the song you're listening to. It samples small bits of audio, creating a digital "fingerprint" of it, and comparing it against a a database on their server to identify the show or movie. In fact, Alphonso's CEO says they have a deal with Shazam, and use their specific technology to do this. But this embedded software can even be listening even when your phone's screen is turned off and it's ostensibly idle.