How does Shazam work? (2022)
cameronmacleod.com
cameronmacleod.com
https://www.wsj.com/video/series/in-depth-features/how-shaza...
Chris (Shazam co-founder)
I wonder if there's a way to build an embeddings model for this kind of thing, such that you can calculate an embedding vector for each new song without needing to fully retrain.
Siamese networks are what you want, two identical pairs of layers (one cached in this case) which act as the fingerprints then then the final layers are doing the similarity matching
People who are highly skilled at this, can be easily stumped. Sure it might workfor artist who are more focused (tailor swift), it might pick out some interesting guest appearances (Eddie Van Halen on Beat It) but when you get multi talented performers who change everything about what do, they don't fit a "model". The most current example would be Andre3000's latest release.
Certainly a well trained model will be able to have incredible accuracy just with vocals alone. It will be able to identify Lady Gaga regardless of whether she is singing a new art pop track or old standard with Tony Bennett.
We could have a debate about the consistency of Gaga or Taylor Swift and profit a motive (and we could go all the way back to composers of the classical period with this).
What about all the people who back pop artists? I dont think picking out Wrecking Crew is gonna be possible (It might but harder) https://en.wikipedia.org/wiki/The_Wrecking_Crew_(music)
I could also point you to Diplo who, as a "producer" is responsible for diverse sounds with his name directly on them and then side projects Like Major Lazer or MIA's paper planes that have his hallmarks but aren't "musicaly" linked. How about the collected work of Richard D. James, I'm no so sure that all the parts fit together outside the whole of them.
Stuart Copland was the drummer for the police, a very distinct and POP sound. Are we going to be able to use ML to take those works and correlate them to his Film scores? How about his opera? Dave Grohl, Phil Colins, Sheila E, more drummers who became singers, what is the context for ML finding those connections (or people).
John Cages 4'33 is gonna be an interesting dilemma.
DO you think the player piano black hole sun, and C.R.E.A.M cover from Westworld are picked up as stylized choices by Ramin Djawadi, and would it link those to the sound track of Game of Thrones?
Even with all the details it's sometimes hard to believe what talented people can do and how diverse their output can be!
but if it is not that would be extremely impressive! determinism/freewill reduces to shazam!?
whats the training data to predict new song titles? heh
check out this reply from claude2:
>predict the next 3 new song titles from artist Taylor Swift
1. Last Dance with You - A reflective ballad about finding closure after a breakup. 2. Never Getting Back Together - A pop tune emphasizing that the same mistakes won't be made twice in a relationship. 3. 22 Was My Prime - A lighthearted look back on her early 20s as carefree years that can't be replicated.
...
Whenever music is mentioned in conjuction with technology, one artist seems to always - in a very literal sense - pop up like a zombie in a B-movie...Taylor Swift. No idea who this person is or what they do but they appear everywhere, all at once.
It feels like a conspiracy.
A noteworthy mention would be that Sony's TrackID did most likely the same thing on their feature phones a few years before Shazam.
Edit: tho for sure, the Philips algorithm was better than either of ours.
And I might be confusing them with another group but I thought, at the time, they were doing some goofy hash of the highest energy Fourier components -- a source of entertainment in our office. ;-)
I think Geoff had the vision and algorithm from the 90s as part of an ISEF project (!?). We had funding in 2001, when we got the real world go-to-your-car-and-get-a-cd and then we identify it ... using the audio signal alone.... demo working.
With a corpus of hundreds of thousands of songs. Positive match in less than 2 seconds.
Sadly, in 2001 there's no market for such whizbang amazing tech.
Shazam only launched one year after that, maybe the problem was in the marketing not the market itself?
I worked on all the Java infrastructure around the recognition cluster (the latter being handcrafted C and assembly, optimised for specific Intel hardware).
The thing that Shazam got right was not just the core recognition tech, but the business processes and supporting systems around it. I remember how much work Chris had to do to convince the 4 major mobile networks in the UK to give Shazam the same 2580 dialing code (the middle 4 buttons, top to bottom, on an early 2000s feature phone).
A major part of the business is the constant sourcing and ingestion of the latest music, in all target markets (think Afrikaans pop in South Africa), deals with pluggers and record labels, etc. Initially, the back catalog was ripped from CD by a huge team of people in a warehouse, on custom workstations.
Just to be clear, it's not turning each song into a hash.
It's turning each song into many hundreds (thousands?) of hashes.
And then you're looking for the greatest number of mostly-consecutive matches of tens (or low hundreds) of hashes from your shorter sample.
Also, I don't think this would be done with training a model today, because you're adding many, many new songs each day, that would necessitate constant retraining. Hashes are still going to be the superior approach, not just for efficiency but for robustness generally.
You wouldn't necessarily need to retrain that frequently. If your model outputs hashes / vectors that can be used for searching, you just need to run inference on your new data as it comes in.
Trendy ("modern") is not necessarily better.
Ultimately this looks the same, but the "hashes" come from a convnet now. But you still are doing some nearest neighbor thing to actually choose the best match.
I imagine this is what 90% of MLEs would do, not sure if it would work better or worse than what Shazaam did. Prior to knowing Shazaam works, I might think this is a pretty hard problem, knowing Shazaam works, I am very confident the approach above would be competitive.
The ml approach is to define a family of data augmentations A, and a network N, such that for some augmentation f, we have N(f(x)) ~= N(x). Then we learn the weights of N, and on real data have N(x')~=N(x).
The denoising approach is to define a set of denoising algorithms D and hash function H, so that H(D(x'))~=H(x). This largely relies on D(x')~=x, which may have real problems.
So the neutral network learns the function we actually need, with the properties we want, where the denoiser is designed for a proxy problem.
But that's not all...
Eventually our noise model needs extending (eg, reverb is a problem): the ML approach adds a new set of augmentations to A. This is fine: it's easy to add new augmentations.
But the denoiser might need some real algorithm work, and hope that there's no bad interaction with other parts of the pipeline, or too much additional compute overhead. (And de-reverb is notoriously hard.)
Generally it's much easier to generate noised pairs from clean input than it is to do the reverse, i.e. go record lots of noised inputs from the wild and match to the original song. So the denoising problem you mention would be tougher still due to covariate shift. I think the features you learn trying to fingerprint the song through noise will probably be a bit more robust, but I don't have a mathematical proof.
Using a model to deconstruct a song like that might enable the ability to recognize someone playing the opening bars of Mr. Brightside on a piano in a loud bar as well as its drunkest patrons.
Nitpick, but Shazam launched in 2002 as a dial-in service that replied with a text-message of the result. The first phone app was for BREW in 2006.
The 2008 date is just when Apple launched the app store; it was not possible for a third party to make an iPhone app before 2008.
In the UK you dialled 2580 from your (non smart) cellphone, it would hang up after a few seconds and you’d get an SMS right away with the ID of the track
It makes more sense if you think of a production like AGT less as the reality show it pretends to be, and more as a promotional reel for labels.
Of course the content they choose to promote is indexed.
> Alphonso's software uses the same technology that Shazam and similar services employ to automatically detect the song you're listening to. It samples small bits of audio, creating a digital "fingerprint" of it, and comparing it against a a database on their server to identify the show or movie. In fact, Alphonso's CEO says they have a deal with Shazam, and use their specific technology to do this. But this embedded software can even be listening even when your phone's screen is turned off and it's ostensibly idle.
It fits into that space of an unusual but comprehensible problem, unlike superficially similar features like recognizing animals or objects in images, which is mostly weird ML magic.
Rather, matching two recordings of the exact same performance (one ingested by Shazam at training time and one ingested by Shazam at run time) is more akin to identifying individuals (facial recognition) than identifying species.
You're dead on that it's pretty difficult if you don't benefit from others, we did a ton of work that in retrospect wasn't necessary. I liked the advanced psychoacoustic model, faithfully implemented in high performant C direct from Zwicker. (Psychoacoustics). To a first approximation, about 10/s model -> pca -> top 16 dim -> VQ and the resulting bytes contain more than 50% of the entropy (!!) Shove all of those in a home grown what-you-now-call-a vector DB, do dozens of range queries, and search for any song common to multiple results. Boom, music recognition. Understandable in retrospect but things like that aren't Everest they're like... multiple unclimbed mountains.
0. And far too early to have any applications. Company existed 2000-2001 \o/
https://patents.google.com/patent/US7853664B1/
https://patents.google.com/patent/US6941275
Very interesting hearing about all of the differing approaches people have taken to solving this problem! Do you have further writings on this topic?
I'd argue that Shazam doesn't have ads, rather it is an ad. You search for the song, then see links to buy it in Apple Music. You'll also see "subscribe to Apple Music" type widgets on just about every screen on the app.
What Soundhound does these days that Shazam doesn't (I think; I haven't actually tried Shazam in a long time) is that it displays lyrics for many songs, and is often able to synchronize those lyrics with where you are in the song.
This is a great post that captures what a spectrogram does, and a must read for people who want to understand how audio fingerprinting works.
There are similar approximate algorithms available for other media as well, so anyone who wishes to understand real world hashing should take their time to study this article.
Every Noise At Once https://everynoise.com
Every Noise at Once - https://news.ycombinator.com/item?id=26668426 - April 2021 (94 comments)
Every Noise at Once - https://news.ycombinator.com/item?id=20585447 - Aug 2019 (82 comments)
Every Noise at Once – an algorithmically-generated scatter-plot of musical genre - https://news.ycombinator.com/item?id=10269685 - Sept 2015 (23 comments)
An algorithmically-generated scatter-plot of musical genres - with samples - https://news.ycombinator.com/item?id=9315499 - April 2015 (3 comments)
shows you similar songs and imo does a pretty good job of it
It's more-or-less identifying melody fragments*, and then just trying to match those up in a sequence. The same way we'll recognize something after 5 or 7 or 10 notes.
I'm pretty sure I've read about other methods for song fingerprinting that rely on things like loudness peaks, where it might work equally well, but that doesn't match how our own brains do it at all. It's pretty cool that this isn't relying on "artifacts" but basically works the same way we do.
* Technically not always melody, but probably is most of the time
Think tapes, wow-and-flutter, speed-up-then-down-all-the-time..
AFAIK fingerprinting is highly time-sensitive (unless cut into ~50ms pieces.. and still not quite).
Last time i looked, the general technique for that - Dynamic Time Warping ? - was prohibitely compute-expensive..
How Shazam Works (2003 Paper) - https://news.ycombinator.com/item?id=33299853 - Oct 2022 (1 comment)
Creating Shazam in Java (2010) - https://news.ycombinator.com/item?id=32530056 - Aug 2022 (36 comments)
Shazam turns 20 - https://news.ycombinator.com/item?id=32520593 - Aug 2022 (227 comments)
How Shazam Works (2015) - https://news.ycombinator.com/item?id=23806142 - July 2020 (7 comments)
Designing an audio adblocker - https://news.ycombinator.com/item?id=18855029 - Jan 2019 (186 comments)
Show HN: A radio/podcast adblocker featuring ML and Shazam-like fingerprinting - https://news.ycombinator.com/item?id=18459058 - Nov 2018 (2 comments)
Show HN: Shazam-like acoustic fingerprinting of continuous audio streams - https://news.ycombinator.com/item?id=15809291 - Nov 2017 (76 comments)
How Shazam Works (2015) - https://news.ycombinator.com/item?id=15350729 - Sept 2017 (13 comments)
Tell HN: Shazam picks up song from my kitchen light - https://news.ycombinator.com/item?id=11593305 - April 2016 (2 comments)
How Shazam works - https://news.ycombinator.com/item?id=9870408 - July 2015 (48 comments)
Patent infringement claim re: “Creating Shazam in Java” blogpost (2010) - https://news.ycombinator.com/item?id=9594480 - May 2015 (18 comments)
The Shazam Effect (2014) - https://news.ycombinator.com/item?id=9593429 - May 2015 (37 comments)
The Shazam Effect - https://news.ycombinator.com/item?id=8634357 - Nov 2014 (34 comments)
Ask HN: Is there an audio search technology that finds exact and similar audio? - https://news.ycombinator.com/item?id=8420141 - Oct 2014 (3 comments)
Source code example of the Shazam algorithm - https://news.ycombinator.com/item?id=5724442 - May 2013 (16 comments)
Creating Shazam in Java - https://news.ycombinator.com/item?id=5723863 - May 2013 (43 comments)
An Industrial-Strength Audio Search Algorithm (Shazam) - https://news.ycombinator.com/item?id=2621103 - June 2011 (4 comments)
Shazam's Search for Songs Creates New Music Jobs - https://news.ycombinator.com/item?id=2215295 - Feb 2011 (1 comment)
How does the music-identifying app Shazam work its magic? - https://news.ycombinator.com/item?id=2214992 - Feb 2011 (2 comments)
Implementing Shazam with Java in a weekend - https://news.ycombinator.com/item?id=1702975 - Sept 2010 (23 comments)
Shazam: not magic after all - https://news.ycombinator.com/item?id=909263 - Oct 2009 (28 comments)
How does the music-identifying app Shazam work its magic? - https://news.ycombinator.com/item?id=893353 - Oct 2009 (16 comments)
I ended up using sox + some Perl to be able to:
- filter out only the frequency of the gun cycling
- identify the peak amplitude to know when a shot started
- use a sliding window of about 30ms to figure out when the shot ended
- once the above was done, I could count how many shots per second and output it
I even posted the code which you can see here: https://www.pbnation.com/showthread.php?t=3216349&highlight=...
I think the historical log of songs that the phone heard launched a little later, on the Pixel 3.
When a "phone" was a dumb device just barely capable of implementing GSM and displaying a clock then this might be worth something as a business, but given where the $0 baseline is, I don't see enough margin to justify competition, I'm surprised even Shazam still makes commercial sense.
Isn’t Shazam owned by Apple now? It doesn’t need to make “financial sense” if it’s a service Apple runs.
Recently, in an effort to simply things, I moved over to Shazam. It’s owned by Apple now, so it’s already built into the iPhone, even without the app. The app allows for saving things a bit easier and I find it to be a lot cleaner than the SoundHound app.
Once upon a time, back in Russia we had a service that captured most of FM radio stations and detected what was playing real-time as you listened from Moscow to some obscure station 2000 km away. Sadly it was all trampled by copyright idiots by about 2013.
But the technology, the capture boxes which were SDRs before SDRs were a thing, hashes, station bosses calling in the night because DJs got too drunk and went rogue live.. oh the memories..
New music is created every day. Can Shazam just buy MP3s at "normal prices" (non commercial) and use them commercially?
Also if they buy music to encode its signature, wouldnt that be very big part of their running costs?
What if someone makes a track and asks 100k USD for it? Will Shazam recognize it? I doubt they want to pay some ridiculous money just to recognize something.
It's like those AI backups that seem to use books for training. Did they pay for those books?
(Screaming electro industrial for those that are curious)
This is different but very similar and contemporaneous.
Much the same technologies though and an interesting research project: https://www.echoprint.me/
And from 1998 Melody based tune retrieval over the World Wide Web https://researchcommons.waikato.ac.nz/handle/10289/1062