Shazam turns 20
apple.com
apple.com
Fun quick related story, about 10 or more years ago there was a back tracking song on a TV show (Scrubs) that I really liked that was only in the Netflix version. It was just an instrumental song with some French sounding words speaking in it so there was no easy way to search for it. However, it was distinct enough that it didn't seem like something made just for the show. It was also pretty quiet and under some talking in the tv show scene. I had posted on reddit asking if anyone knew it, and never got any responses. I searched all over the web, but no source had the track details. It drove me crazy every time I would hear the song in re-watching the show, and I still could not track it down every few years when I tried again. Back then, Shazam had no cataloging of it so it wasn't in there either yet. However, when re-watching it a few years back again, I tried Shazam again and to my surprise it finally worked. I was blown away that Shazam was finally able to solve this 10+ year mystery. It was one of the coolest feelings every to scratch that itch finding this rare French song and hearing it in full. It was truly magical.
EDIT: Oh sorry, I didn't think anyone would actually care about the song itself lol It was called "Sans Hésitation" by the French-Canadian band "Chapeaumelon". https://www.youtube.com/watch?v=Ju4d3YQhByU - It's also interesting cause now the song does in the episode in tv music database sites. Very cool.
Personally, my main usage of Shazam is for identifying vaporwave samples. Often all you have to do is throw the song in Audacity, tweak the speed a bit, and Shazam it.
I ask because I like to create bootlegs (basically homebrew remixes, these are substantial re-imaginings of the original track) and would like to put them on Spotify, but am worried about copyright issues and how that might affect posting original music.
1. https://www.whosampled.com/Washed-Out/Feel-It-All-Around/ 2. https://www.whosampled.com/Macintosh-Plus/%E3%83%AA%E3%82%B5... 3. https://www.whosampled.com/song-tag/Vaporwave/
Yeah - I think it's magic as well.
Other thoughts: I used it back in the UK when it launched, and the first track I ever used it on dialling (2580 - the numbers down the middle of your keypad) was also a French track (MC Solaar – La Vie Est Belle)
I always felt they missed a trick, just identifying music (and then trying to sell you stuff). Surely they could have used the same tech to seamlessly mix all music together. (i.e. take the sequences within tracks they find hard to differentiate, and then use these points to allow two tracks to be mixed together). What's the minimum number of tracks it would say take to seamlessly mix from Megadeth to Mozart?
I believe it's very sensitive to changes in timing, so it doesn't work on live performances etc.
(based on reading I did 13 years ago before an interview at Shazam, which to this day still remains my worst interview performance)
What we do is calculate the acoustic fingerprint of every uploaded content and compare/check for duplicates (only authorized staff can upload, but this still helps a bunch with user errors and in cases where you need to reupload a track). Then we compare the fingerprints, using this[1] approach, so we can fine-tune the similarity based on our needs.
In our case it's been very effective. Yes, live versions are treated as different ones (which is exactly what we need in our case, so it's a feature for us), but mechanical differences between tracks (volume, slight distortions from codec, different compression levels or remasters, or track being cut differently) are just ignored.
If you ever want/need audio fingerprinting, I can warmly recommend it.
[0] Music streaming service optimized for cafes, restaurants and other venues - https://musicbox.com.hr/ [1] https://groups.google.com/forum/#!msg/acoustid/Uq_ASjaq3bw/k...
I think you're talking about a live recording vs a studio recording? But what I think zelos was talking about was "someone is currently playing music live, what is it?", which is a lot harder because you need to recognize the essence of a song and not the essence of a recording of a song.
One of these might match random points in many songs, but a far smaller subset of these will have the same three in the same sequence.
As far as I can tell these operate on audio, not symbolic music.
I noodled around with this idea in my free time a few years ago, got absolutely nowhere really usable with it (I probably put in a couple hundred hours).
I knew I was limited by my dataset (small), code quality (terrible) and understanding of musical theory (virtually nil).
Maybe I'll pick up that idea again - even doing beat matching would be kind of neat.
They must have loads of data on songs people actually want to know yet never really managed to turn themselves into anything more sophisticated.
It does a Fourier analysis of sections of the song, and puts the results in a database. A Fourier analysis yields what frequencies make up a waveform along with their amplitudes, so it is very compact.
It's not the analysis that is compact, but the fingerprint derived from it.
...who by the way holds a PhD from Stanford...
I was wrong, of course, Shazam really did live up to its hype. I think it's interesting that the someone knows about how a technology works the more sceptical they are of what it is capable of.
Which I always find to be simultaneously simple and obvious as well as total magic.
Way back then, it was doing everything you describe, but over low quality band limited telephone lines.
And while it took a few iterations (for me, from palm pilot to blackberry as a teenager, then eventually moving to iPhone after a few too many painful Blackberry upgrades - still missing that unified inbox though, as is everyone else I know who had a BB of that era... and frankly missing a great physical keyboard on a phone, too) I still am impressed on a daily basis that I do indeed have the device in my pocket that 12 year old me dreamed of.
They eventually shut it down :( https://slate.com/technology/2013/05/google-sms-search-shutd...
Digital phone calls today are way better.
It looks like others shared the paper: https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf
It's short but very cool. I read it a while ago and honestly can't pretend I fully grokked everything, but my understanding was that you can't just use a Fourier transformation alone. Noise would basically make this impossible.
So what I'd consider the key insight is that they compressed songs down to "fingerprints". IIRC they noticed that songs, even in noisy environments, preserved certain bits of information. Particularly, they could look at the spectrogram and see peaks of amplitude in the tapestry. They essentially set some radius and scanned the spectrogram. In a given radius, only the largest amplitude value in time and frequency would be preserved. So you've reduce a 3MB song to several bits.
This would be good enough for small databases (I think). But it's intractable for anything practical. So they built hashes out of these fingerprints using pairs of the preserved peak bits. They would choose a certain peak (called the anchor point), record its time offset from the start of the song, and then form pairs with other nearby peaks, saving the pairs of frequencies (but discarding e.g their amplitudes). So for each of these anchor points, you would get a 64 bit value: 32 bits for the time offset and track ID and 32 bits of frequency-pairs.
When you wanted to look up a song, they would fingerprint your snippet into multiple 32bit hashes and compare them against the frequency-pair hashes in the database. If a song was a good match, then you would see that your snippet matched against multiple hashes from that song, and specifically they matched linearly over time (I'm struggling to explain this bit but it's visually obvious if you look at Figure 3 in the paper).
I probably got some of this wrong, but I hope it's a helpful summary of the paper. I remember struggling to understand parts of it, so please let me know if anything I said is egregiously wrong!
Wow blew my mind was when Google introduced 'hum and we'll recognize the song for you' in Google assistant: https://www.google.com/amp/s/blog.google/products/search/hum...
It works so well even with my shitty humming - even my girlfriend can't recognize what the song is but Google can. It doesn't even have the same signature as the original audio file, just similar hums in a noisy environment and it still works. Black magic fuckery.
Totally agree though. It is something that opened my mind to thinking of a way to solve that problem in a way that actually works. Shazam definitely looked like magic the first time I saw it work.
To me, this highlights how hashing is the closest thing programmers have to magic.
Their announcement actually made me roll my eyes a bit, as Soundhound had that functionality nearly a decade before. I had both SH and Shazam installed on my old phone for these usecases - now Shazam is baked into Siri so I don’t even have the app itself installed.
There's also a bunch of other options to trigger Shazam, main way I use it is from the Control Center: https://support.apple.com/en-us/HT210331
Soundhound is what had humming “support” explicitly in its product description, and it worked pretty well from what I remember. It’s been long enough though that I may only be remembering the times it worked.
You aren't giving it enough credit. The algorithm uses just a few seconds from any part of the song, and has to deal with phone audio quality and often background noise. I mean, you can be in a bar with all that jabber and hold up the phone and it could pick out the song. The app on the phone does the preprocessing to the audio before it is sent to the server that does the matching ... using the comparatively miserable power of a 2001 era cell phone.
And Dall-E 2 is just doing fuzzy hashing of images with text keys.
Shazam continues to amaze me because it "just works", and still feels more magical to me than most of the AI out there since it directly solve a major problem I didn't even think was solvable "what is this song!!?"
Basically, it records audio chopping it up into small segments and throwing them through a FFT. Then it takes that, and thinking of the data like a greyscale spectrograph image, runs it through a quantization filter that helps reject some noise, then converts that to locality sensitive hashes that are sent to the server. So basically FFT, filter, hash, lookup.
On the topic of background music, tons of original background music copies/imitates famous stuff. Sometimes it's "I wanted the sound of X but couldn't afford it", but there are some in-jokes in there too. Wish I could remember some examples.
Avery had gotten his PhD from CCRMA at Stanford under Julius O Smith. His PhD had been on the topic of automatically ("blind") recovery of individual vocal/instrument tracks from a final mix. From there he joined a startup, Chromatic Research, where I was also at. He created lots of code and patented some algorithms related to resampling and MIDI synthesis, stuff like that. Avery was (and is) a super nice guy, humble but incredibly smart -- he could work not only with the high level mathematics but was also equally excited to fine tune assembly code.
After Chromatic folded, Avery had been struggling to get his own startup off the ground. About the same time, the Shazam guys had the idea for the product but didn't know how to create the algorithm. They approached Smith at CCRMA looking for someone capable of creating something that worked. Smith suggested they try Avery Wang.
At first Avery said, "Hmm, that seems difficult, but let me think about it." Within a week he had a demo running on a few thousand songs he had gathered from CDs. I'm sure a lot of refinements went into it after that, but the core idea took him a weekend.
[ All factual mistakes above are due to me and 20 year old memories -- if I misrepresented something it certainly wasn't due to Avery telling me something that wasn't true ]
EDIT: A blurb about Avery https://www.seti.org/avery-wang
Incredible! Curious to know what exactly happened backend after it listened to the audio, and what hardware it ran on.
AFAIR, we never had CDMA in the UK, so what Verizon et al. were using is irrelevant.
It’s the same as recognizing objects in a 256x256px image.
Try resampling a song from 44kHz to 4kHz and you’ll still have no trouble recognizing it.
A 6 string guitar for instance goes from E2 (82 Hz) to E6 (1318 Hz) for 24 fret electric (classical guitars typically have 19 frets and go to B5 (988 Hz) and acoustic guitars have 20 or 21 frets so go to somewhere in between).
Popular singers with high notes are Mariah Carey who goes up to G7 (3136 Hz), Christine Aguilera who reaches C#7 (2217 Hz), and Prince who could hit B6 (1975 Hz) [1].
Plain old analog telephones and, I believe, early cell phones had a voice band of 300-3300 Hz.
They would have no trouble with most of the notes in the upper parts of the aforementioned ranges, except for the top 5 notes of a piano. As you note you'd change the timbre, but you'd still have the right notes.
Low notes might be a problem though. If you lose everything below 300 Hz that would cut out most of the left hand on a large majority of piano parts. On guitar it would cut all the notes that cannot be played on the first strings except for one.
That would change the notes. You'd lose the fundamental of a lot of notes just leaving the overtones, so it would look like the musicians played a higher note.
My guess is that when they were processing the song database to generate the hashes they put the songs through a bandpass filter the was smaller than the frequency range of the most limited device they supported listening on. Then when listening on any other device they could filter it down to that so those hashes would work.
[1] https://www.concerthotels.com/worlds-greatest-vocal-ranges
I don't really understand why phone call fidelity hasn't improved since then. Sometimes it seems like it's even worse!
No? I guess the hypothesis that audio fidelity is crucial was wrong.
It wasn't just a case of developing an algorithm that could in theory be used to match an audio signal against all the world's pop songs. They presumably also had to get hold of a substantial number of those songs, fingerprint them, and roll out the search robustly against generally very poor audio hardware using simple telephony services at (for the time) quite considerable scale. They did it very quickly, it worked super well from launch, and it's been running continuously ever since.
I've read the paper about the method, but I would love to know more about the original development and deployment.
I used Shazam in London when it launched. I left the UK in 2003 so I reckon it was literally just as it launched. Presumably I saw billboards or a Wired article or something?
(1) https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf
And most of all: no ML involved! All hail the heuristics!
To this day, Shazam still has that aesthetic of "press big button, do magic".
Only Google has managed to top Shazam in blowing my mind, and only ~recently, by making this whole process happen completely offline and continuously in the background on a phone. It's not as broad but still incredible. Google's paper: https://arxiv.org/abs/1711.10958
What I do is to imagine myself finding a smartphone in elementary school (90s kid). These are a few things that would blow my mind:
- Having a digital global map, with multitouch, that can show me where I am in that map. I can search anything and find reviews from virtually anywhere in the world. I can zoom and see my actual house. I can use street view.
- I have access to any song I want.
- The phone can listen to a song and it can tell me the name of it (then I can listen to it again)
- I can play video games with much better graphics than my N64
- I can watch movies and TV in there.
- I can video call
- I have a digital assistant
- I can find any answer online
- I can buy anything online
- In the future all this technology is not just for the rich, virtually anyone can buy a smartphone.
And - you can play N64 games on a rectangle you can carry in your pocket.
Fastest is simply turning on “Shazam on app start” in the settings and then starting the apps. However if your phone is locked then indeed what parent commenter said is quickest.
It really felt like pure magic!
I have an image in my mind of my boss at the time going around the office asking if anyone was interested in talking to this thing called Shazam. I've long wondered if I imagined it. I certainly didn't act on it.
I remember (not much later than this) interviewing at a place where the product was intended to be "an automated assistant that listens to your phone call and pipes supporting information to your computer as you speak". Obviously I gave them a wide berth. It's funny to think about the "gap" in magic - Shazam seems magical but totally worked, this other idea seemed magical and, at the time, totally was.
> Role: Senior software engineer - Low Level Device , Distributed Communications Role mission: To ensure that Shazam's subsystems are integrated and interface effectively and efficiently with external partners' systems/hosting environments, yielding available, robust and scalable full offerings. Key Performance Areas: 1. Design real time software using standard techniques and protocols, to be scalable, maintainable and robust 2. Manage & collaborate within and between team(s) 3. Implement quality software solutions within budget 4. Ensures that design and implementation of software is of high quality 5. Ensures that all deliverables are documented Required Skills/Capabilities <B7> Knowledge of interfacing peripheral and devices to Linux <B7> Knowledge of Linux device drivers a plus. <B7> Distributed messaging techniques and protocols, eg: PVM, MPI <B7> Ability to grasp and work with abstract concepts <B7> Familiar with current software engineering methodologies e.g. RUP, XP <B7> Understands and is able to manage quality assurance e.g., module tests, code review Required Knowledge/Previous Key Experience <B7> At least 4 years of full-time software engineering within a team of at least 3 sofware engineers. <B7> Must have been involved in all phases of the software cycle from requirements engineering to launch. <B7> Must have developed low level device or communications software <B7> Experience with Computer telephony a big plus <B7> Experience with a high-growth startup environment a plus Ideal Qualifications Ideally University degree in Computer Science (alternatively at least 4 years of proven software engineering experience). Please forward your CV/resume', with cover e-mail, including full details of your earnings expectations, to recruit <at> shazamteam.com
Wrote about it too: https://umaar.com/blog/lessons-learned-from-working-at-shaza...
[1]: https://www.apple.com/newsroom/2018/09/apple-acquires-shazam...
The site seems no more but I found a Lifehacker post about it: https://lifehacker.com/find-the-name-of-a-song-by-tapping-14...
A decade later I discovered Shazam, and even today, more than a decade after that, Shazam still has a place on my home screen, quickly within reach, helping me discover hundreds of great artists and songs overheard from as many different places. The magic of the experience, and the appreciation for the technology, stem from the memory of that moment in the mid-nineties when I stood under a speaker listening to a song that I might never hear again.
I asked how it went, and he typed something like "du du duu duu du du duu du, du du duu duu du du duu du" and within 10 seconds I replied "oh, Tom's diner by Suzanne Vega?" After a few moments he replied "yes! how the hell?!"
Anyway, Shazam is great when out and about and I hear something I like. Clubs and other loud venues provide a challenge, but covering the mic usually does the trick.
I'd love to read some more details about how such fingerprinting works. I'm sure there are lots of interesting details on how it deals with recording noise and such.
There's more to Shazam than that but Fourier transforms gets rid of the noise. I ported a FFT to Java back in the days and it was, IIRC, not even 100 lines of code. Amazing algorithm. I used it to record engine noise under acceleration and then derive power/torque curve of my car (it took into account the number of cylinders): drive the car several times, both ways, on a street, record the noise. Apply the FFT. Input the rims size / gear ratio etc. And I'd end up with about the exact same plot as the official one from the car manufacturer.
Noise simply disappears with a FFT.
A more concerning issue is harmonics.
#1037 +(972)- [X]
<Mass> hey does anyone know what the song name is by Frankie Lymon that goes Uhhh uhh uh uh uhahhhhh uhh uhh uhI also love that it just shows up on my lock screen.
Recently for whatever reason I was listening to the twist/ cover "Stacy's Dad" and Now Playing recognised it as the rather more famous original Fountain's of Wayne "Stacy's Mom". So yeah, it doesn't know everything. It also doesn't recognise lots of obscure stuff I own that's like B-sides or special editions that never saw radio play, or bands that my friends were in (but everybody I know has read both Steve Albini's "Some of your friends are probably already this fucked" and the KLF's "The Manual" and so none of them signed a recording contract and thus you've never heard of them) but I've never had a situation where I heard something I liked at like a bar or somewhere and Now Playing didn't know what it is.
https://open.spotify.com/playlist/2Z69DiD4sGs5aHt4OXE3fU?si=...
There are other apps such as SoundHound, and even some virtual assistants have the same feature.
Even more obscure is the list of all songs you ID is maintained. Launch the “iTunes Store” app (why is this still even a thing) then tap the upper right hamburger menu. Then select the middle “Siri” tab. It’s kind of neat. I use this feature when I’m out and about in restaurants, stores, etc.
Again no need to install any app.
One used to be able to sing the songs as well which does not seem to work anymore. Either that or my vocals are gone.
I expose the extent of my confusion here: https://twitter.com/walkaboutrick/status/1560713609948250113...
Source: I watched the Shazam founder awkwardly pitch Danny Boyle at a director meet-and-greet while Danny tried his best to avoid them.
CLI app not Python notebook, btw.
edit to add: I only know this so well because, when I switched to iPhone in 2016, I was REALLY confused about how you Shazam something. I couldn't find it in the App Store. I tried to use the Google app, since that's how I did the equivalent thing on Android, but it didn't support it. Finally I figured out that it's just built into Siri.
ERGO I am surprised that they put out a PR about it.
That’s probably because no such history exists, Excel was always a Microsoft product (even if they aren’t the inventors of spreadsheets).