Building a Music Recommender with Deep Learning
mattmurray.net
mattmurray.net
So, I believe that this is actually supervised learning, as the author is training a classifier on preexisting labels (the genres).
I believe that unsupervised learning would not make use of a target variable at all. If the network architecture terminated at the fully connected layer, and then propagated that layer backwards to reconstruct the input (something like Contrastive Divergence), that would be an unsupervised method.
I understand this is an educational project, but nevertheless it's published, hence open for critics ;)
Edit: small style corrections.
The first step to analyze this is to make a confusion matrix, [1]. It would be nice if the article included it.
The results are good though! Good work! :D
[1] http://twistedsifter.com/2013/01/hidden-images-embedded-into...
Yes, and thus the reason why the classifier was so good at recognizing trance...it's one of the few genres that locks in at around 144bpm.
[1] http://journals.plos.org/plosone/article?id=10.1371/journal....
The spectrogram is just a series of FFTs taken over time; encoding it as a bitmap doesn't really change this, aside from precision issues. Any other representation of the audio is derived from either the original time-domain signal or the FFT.
Indeed, humans can't reliably map raw waveforms or spectrograms to intuitive musical phenomena. But a CNN should be able to derive meaningful features from these basic representations on its own.
The history of recording industry "Genres" has close ties to cultural segregation. Pandora's Music Genome approach is optimized to break the genre barrier.
It'd be interesting to see how many "Down tempo" songs shared characteristics with "R&B", for example. I think the Author's approach could still be applied.
Perhaps. But of course, this is likely to put the user literally into an "echo chamber" :)
Is R.E.M. "Alternative", "Rock", "College"? Maybe you consider an album like "Reckoning" from R.E.M. "Rock" but then it includes a track like "Rockville" that is perhaps "Country"?
Genre makes sense for "Soundtrack" or perhaps "Classical"? But beyond that it's just mental gymnastics.
And given how fondness for music is qualitative, I've always been suspect of any sort of algorithm that tries to recommend music based on fast-Fourier-transforms. Maybe AI isn't for everything....
http://webcache.googleusercontent.com/search?q=cache:http://...
http://benanne.github.io/2014/08/05/spotify-cnns.html (Recommending music on Spotify with deep learning) uses CNNs trained on spectrograms + similarity data from collaborative-filtering to predict per-song vectors.
There are a number of interesting directions you could go with that data set. One interesting possibility is to make a convolutional autoencoder, then use that to apply "deep dreaming" filters to music. Another interesting evolution would be to handle the frequency dimension using a 1D convolution, and run a RNN on top of that to deal with time.
He's taking 185000 samples, and finding similar "looking" samples elsewhere in other songs, and then making recommendations based on that. I don't see what that could possibly have to do with genre labels, unless we're under the assumption that finding a match between a Drum & Bass song and one that seems similar with a tag of Trance is somehow a bad match? (which very well could be the case, but seems like a big assumption to make off the bat)
Are these recommendations silo'd to the current genre or are they allowed to span genres?
https://static.googleusercontent.com/media/research.google.c...
Within one arbitrary song (Inifinite Jukebox - no longer working?): http://labs.echonest.com/Uploader/index.html
https://www.reddit.com/r/infinitejukebox/comments/4cmr4f/met...
My first thought was to wonder how a LSTM would do. Once might think it would be a better representation for music? There's some models which use convolutional layers along with a LSTM for video representation (eg [1]) and it would be interesting to see if convolutions are useful for capturing similar themes of music.
I wonder if one could build a music embedding (word2vec style) and use similarities in the embedding space as recommendations? The obvious objective function would be skip-gram, but there might be more interesting objectives there too.
[1] https://github.com/loliverhennigh/Convolutional-LSTM-in-Tens...
I completely agree LSTM would be useful as it would by default require a different representation. I think most commenters agree this representation is overly simplistic. Amazed it works as well as it does!
https://www.facebook.com/photo.php?fbid=10154605399547143
> Don't you guys realize that putting everything from Monteverdi to Bach, Mozart, Beethoven, Brahms, Moussorgsky, Stravinsky, and Bernstein in the same "Classical" bucket makes no sense?
> (Particularly when you have ultra fine-grained categories for popular music!)
Any comments about that?
However, 2D conv+maxpool is an image processing technique that gets you translation invariance. Fine for the time dimension of the spectrogram, but rather dubious for the frequency axis. Surely you'd want to distinguish if some feature happens at a high or low frequency?
MFCCs[1] are exactly that, a type of convolution along the frequency axis of a Fourier transform, and are highly apt features for music classification tasks.
It makes sense if you think of timbre as a time-varying relationship between the harmonics of a single pitch; translation invariance along the frequency axis can tell you that you there are partials typical e.g. of a guitar or of a flute, without caring what particular pitch those instruments are playing. And timbre is a bigger source of variety in popular music than e.g. the particular notes used.
Why not checking which are the top 3 most played songs by other users who are the 1000 users who have the most similarity with the current user, and then recommend the current user the most played songs from the 1000 similar users that the current user has not listened to yet.
As far as I can see this would be superior to any existing A.I. recommendation algorithm.
My 1 min effort description would be biased towards popular songs, but you can easily change that by selecting songs that are not popular, but that occupy a lot of playtime with a user.
This an interesting approach, but the objective is similar to most recommendation engines: "Find me something similar to something I like". Sometimes that's a good requirement (e.g. when trying to queue up the next song in a playlist, it's good to have some similarity to the song you're currently listening to). However, when trying to discover new music it's generally a bad approach; since (depending how the requirement is tackled) you'll get recommendations that tend towards some median; i.e.:
- Other songs by the same artist
- Songs by artists who have collaborated with the current artist
- Popular songs (i.e. if almost everyone has a Beetles album in their playlist, getting "people who bought this also bought" recommendations for anything would list Beetles, since technically that's true; it's just uninteresting.
- Songs in the same genre
- Songs with a similar sound / structure
i.e. it tends to list things which you're likely to be aware of anyway. Also this means you'll get lots of songs with little variety between them; making your playlists monotonous.What I'd be really interested in seeing was an engine which finds things on the peripheral; i.e. figures out the things that are likely to appeal to you because of the more unique things you're interested in; or the popular things that you dislike. That way you're likely to get a more eclectic mix of suggestions, and broaden your musical awareness. This would likely produce a lot more false positives initially, as it's expanding your taste range rather than narrowing in on some "ideal" average, so may stray into unknowns; but once you've heard and rated something in this new area, that data can quickly feedback into the algorithm and thus you learn of things you'd previously never have discovered.
But as of, 3-6 months ago those daily mixes started putting some really interesting new songs that I wouldn't find otherwise. Sometimes it seems to go back to that "safe zone" but it's been such a much better experience I have been telling all my friends to try it.
I really would like to know more about their process to improve the recommendation system.
This does make a lot more sense than analyzing the audio of the music IMO. For example youtube does this okay and if you look for a Mazzy Star song after watching Ricky and Morty (a tv show), it will recommend other Ricky and Morty soundtracks even if the style is completely different. This isn't something you can predict with just audio data.
I've been learning recommendation engines by looking at peoples' Steam games libraries.
One feature of the data set is that many, many people own multiple versions of Counter-Strike as well as Team Fortress 2. So "a high number people who bought [almost any game] also bought Counter-Strike: Global Operations" is a recurring problem with a naive recommender.
What I've been learning how to do is weight recommendations by how 'surprising' they are, for want of a more accurate term. If 80% of people who own Game A also own Game B, but only 5% of the total population owns Game B, then we should upweight that relationship.
[1] https://books.google.nl/books?id=_AfABAAAQBAJ&pg=PA258&lpg=P...
Reggae as "genre" itself is also quite varied in what goes under its label. There are also other factors that play a big weight on how good matches they are to reference material. Producer and decade make a huge difference but also what's known as "riddim" name should give clues.
Great suggestion / I guess this leads to the idea of needing a meta recommendation engine; i.e. some way to decide what recommendation engine best works for you; selecting from one that follows lyrical themes, another that discovers "out there" content, one for similar content, etc.
If however I want recommendations for new Metal music, and my previous selection was Metallica then you play me some Megadeth, I am going to hate it and not be interested in it at all!
I've completely given up on goodreads for recommendations and just google "best <insert genre> books 2017" now and usually can find some good lists.