Transformers in music recommendation
research.google
research.google
https://cosine.club/ is the closest I've seen to the ideal branching system. My understanding is that it uses vector embeddings to search for songs that are similar in sound, and it works shockingly well for that purpose. However, it has a limited song database. Also see https://everynoise.com/, which is no longer updated. These use vector embeddings in similar ways, but the exploration experience is controlled by the user, not by a list-generating ranking model. I definitely think that AI-tech is the future of music recommendation, but I would prefer to see more research by large companies in to these user-driven systems, instead of the 'similar autosuggested list', which is, by its very nature, only ever 'good enough.'
The reason this simple mechanism works so well is that it gets rid of personal biases and instead taps into a community of listeners listening to the same stuff. Confirmation bias is the core issue here. I don't want confirmation bias. I want my biases challenged with new things. Not randomly new but based on what others are listening to that listen to similar things. And not just randomly based on everything I listen to but on specific things that I'm playing.
Vector similarity of artists could be an interesting angle. But it would probably risk pulling out a lot of cover bands and imitators. You want stuff that is close but not too close.
But you are right that none of this stuff is perfect.
I have a broader concern for how new artists are supposed to get discovered without a promotion engine behind them. Yes it's always been hard to get started, but the distribution of attention has really become much more top heavy in recent years. I know one guy who played Wembley stadium and still couldn't give up his day job which he was sure he would have been able to do following a gig of that size in the 90s. Yeah so he had a good number of monthly listeners, but it illustrates how the distribution has changed.
Plenty of people on the long tail deserve to be discovered, and use of AI to recommend music - in place of collaborative filtering - really has the potential to fix that.
PS. We were talking monthly listeners weren't we, so you'll be excited to know that fractional fans exist already ;-)
Artists struggling to make a living on the back of a single success is, if anything, a product of the longer tails of music being a catered to. The gains are much more spread out now.
It's maybe a niche argument but I'd suggest looking at the one hit wonders of today vs yesterday: https://en.wikipedia.org/wiki/List_of_one-hit_wonders_in_the...
imo the one hit wonders of yesterday were fairly significant hits. The one hit wonders of the 2010s are vastly more ephemeral in my personal opinion. Probably mostly driven by the fact that they used to be conveyed by pop radio and now I don't hear pop radio EVER. But I also have some doubts that most of these 2010s songs will be able to carry a band forward like the one hits of the 90s.
however much of my personal discovery is based on trying to understand the history of groups that i piece together from wikipedia and reading about who the artists were.
I want is recommendations based of some sort of in-depth knowledge graph that traces personnel hopping between bands, which other artists worked in the same scene, who they public acknowledge as an influence, etc.
it would be great to uncover things like "hey, did you know that all these songs you like had the same producer? maybe you should dig into other things that this guy produced" or "this artist you loved was really into a performer from a completely different genre -- maybe you should check it to see the influences that they had"
Any song has many facets: melody, key, rhythm, dynamicity, voice of the artist, lyrics.
I can see how a music browser of the future (rather than an automatic recommender) would be equipped with many different knobs to turn and tweak each of these dimension's weight (as they are going into a similarity calculation) separately, to give the user control.
It requires the music to be already present, however, so not ideal for finding new music.
Considering the pool of music on the radio when I was a kid, that's a reasonable hit rate in my mind. I'm not sure I would cope with an influx of music larger than that.
Genre, tempo, key, vocalist sound, instruments, and so on. These might all be relevant in different recommendations, at different times, in some particular order depending on the user. The music-content in effect only serves to align tracks along lines in the embedding space.
Algorithmic playlists I've not found useful. The Apple Music "create a station from this song" feature is more or less broken imo -- i get so much of the same same same stuff
Apples recommendations are so bad. It just goes right back to the same top 100 songs.
Apple Music is notably better – and has the benefit of not funneling your money to the likes of Rogan – and the recommendations will be fairly good within a genre but it does overweight your library a bit (I wish it had a “I’m looking for something new” / “familiar” toggle).
I am curious what Rdio did differently as I had a very good success rate with their suggestions and it seems unlikely that there was some secret sauce nobody else has been able to figure out.
Have you tried the “Discovery Station” on Apple Music? It’s supposed to play only music new to you. It’s fairly new and was introduced in summer 2023.
Related question: I wonder if identical twins are good at recommending each other music
I doubt that any recommendation system is capable of providing meaningful results in absence of the "awareness" about the actual content (be it music, books, movies or anything else) of what it's meant to recommend.
It's like a deaf DJ that uses the charts data to decide what to play, guessing and incorporating listeners' profiles/wishes. It's better than a deaf DJ who just picks whatever's popular without any context (or going by genre only), but it's not exactly what one looks forward to when looking for a recommendation.
That's the thing with transformers, right? It doesn't actually "know" anything about its inputs.
The embeddings are learned (initialized to random).
Pandora was much smarter, but seemed to run out of songs instantly.
They took all user generated playlists and projected the songs into vectors where songs that appear together on playlists are closer and songs that appear less often are farther.
It’s likely changed a lot since then, but it seemed like a pretty straightforward clustering system at the time.
Or, more bluntly: you aren't going to mate with a For You page, so it doesn't have the same evolutionary cheat code to your preferences as other people have.
This is the same way YT/TikTok does it btw. Co-occurrence is king in recommender systems in production. It's extremely cheap to calculate and by far the most effective method.
This is not really important if you have a lot of user behavior data and/or playlists for each song. But if you have a niche song that few people of listened to, collaborative filtering based recommendations aren't going to be good.
Real semantic embeddings (which can then be part of the input to the recommendation model) can be trained using self-supervision, e.g. an auto encoder or a seperate "next audio token" predicting transformer.
It uses data about overlap in listenership between different artists to determine which artists are related to which others and how. The artists serve the same role as words in sentences.
My experience is that the best music is found randomly. I like so much, I don't even know what I really like. Even what I like is always changing. I need to listen to ton of random things I don't like and I will find a small amount of gems. The absolute gold though is finding songs I didn't even know I would like.
The algorithmic version of sifting through records at a record store for a music lover is random. Random with an easy way to play the next song.
All these recommendation systems are just Satie's musique d'ameublement generators for non-music lovers. Furniture music generators, music to play during a dinner to create a background atmosphere for that activity.
Other times specific song or music genre is relevant to me because of a moment in real life or from a movie.
Most of the reasons people like music, or fictional movies and books, is personal, emotional, subjective, and difficult to articulate. You wouldn't know what data to collect. You're better off just asking them to rate song, movies, or novels out of ten. You can then compare their ratings with other people's, and what you'll find is there are clusters of people who rate things similarly (and others who rate things differently), and that the ratings they give overall somehow capture their feelings about whatever they listened to, watched, or read. (Source: I developed a movie recommendation system which predicted ratings reasonably accurately.)
Of course, if you just have sequences of user actions, like in the article, your recommendations won't be anywhere near as accurate.
Years of experience have proven that you can get quite far with pure collaborative filtering—no user features, no content features. It's a very hard baseline to beat. A similar principle applies to language modeling: from word2vec to transformers, language models never rely on any additional information about what a token "means," only how the tokens relate to each other.
I would like to be able to run this model myself and have a pristine and unbiased output of suggestions
If I try to play any music from a historical genre, it's only about 3 or 4 autoplays before it's queued exclusively contemporary artists, usually performing a cheap pastiche of the original style. It's honestly made the algorithm unusable, to the point that I built a CLI tool that lets me get recommendations from Claude conversationally, and adds them to my queue via api. It's limited by Claude's relatively shallow ability to retrieve from the vast library on these streaming services, but it's still better than the alternative.
Hoping someone makes a model specifically for conversational music DJing, it's really pretty magical when it's working well.
> How do commercial considerations impact recommendations?
> [...] In some cases, commercial considerations, such as the cost of content or whether we can monetize it, may influence our recommendations. For example, Discovery Mode gives artists and labels the opportunity to identify songs that are a priority for them, and our system will add that signal to the algorithms that determine the content of personalized listening sessions. When an artist or label turns on Discovery Mode for a song, Spotify charges a commission on streams of that song in areas of the platform where Discovery Mode is active.
So Spotify's incentivized to coerce listening behavior towards contemporary artists that vaguely match your tastes, so they can collect the commission. This explains why it's essentially impossible to keep the algorithm in a historical era or genre -- even if well defined, and seeded with a playlist full of songs that fit the definition. It also explains why the "shuffle" button now defaults to "smart shuffle" so they can insert "recommended" (read: commission-generating) songs into your playlist.
[0]: https://www.spotify.com/ca-en/safetyandprivacy/understanding...
(i work in advertising and we would never be allowed to introduce sponsored content into an organic stream like this without labeling)
Here are two that might fit what you're looking for:
'90s K-pop: https://open.spotify.com/playlist/6mnmq7HC68SVXcW710LsG0?si=...
'00s minimal techno: https://open.spotify.com/playlist/6mnmq7HC68SVXcW710LsG0?si=...
There are sites to convert from spotify to another service if you don't have it.
Usually I find them _by accident_ while browsing public playlists on Spotify
A transformer may also be larger than their baseline, but you still need to justify how those parameters are allocated.
My favorite band (vulfpeck, and more recently jack's solo stuff) often branch out into different genres, and it's a bit of whiplash when it goes to another song just because the artists are similar / appear together in other places.
> Plex Media Server uses a sophisticated neural network to analyze each track in the music library, cataloging a wide variety of characteristics of the track. Think of it as things like female vs male, vocals vs not, sad, happy, rock, rap, etc. All these various characteristic constitute a “Musical Universe” and the server is determining where that particular track exists within it.
> For the math-savvy, the Musical Universe consists of points in N-dimensional space. But what’s important is that this allows us to see how “close” anything in your library is from anything else, where distance is based on a large number of sonic elements in the audio.
I haven't tried it so can't speak to its effectiveness.
The problem with classification is that what makes a genre is not uniform. Some genres are defined by the way people sing, other genres are defined by the singers language, other are purely about the instrumentation or rhythms used, yet others mostly about the sounds and notes used etc.
But there are things like tempograms, tonnetz (tonal centroid features), chromagrams, spectral flatness/contrast/roloff, laplacian segmentation etc. And I guess feeding these into some neural net might give you interesting results.
Algos are useful for finding and playing music but they are just a tool. The human touch is needed to create the best musical experiences. That's why we need DJs now more than ever.
To experience the best music, for you, you're going to have to do a bit of work to search out new sounds and old sounds. So go digging and create awesome playlists, for you, and share them: they'll probably be better than algo-created playlists.
Algorithmic playlist systems are great but I believe we should counteract their prevalence by also supporting human efforts: by also listening to online radio stations and DJ mixes, supporting music journalism and publications (as well as buying physical media and merch, going to gigs, of course!)
I don't understand exactly what this Google paper is describing but it sounds like it's going to anticipate whether I want to listen to up beat or down beat music. What a very googleish thing to do! - this doesn't help me at all. Also I'm turned off by it's opening statement:
>We present a music recommendation ranking system that uses Transformer models to better understand the sequential nature of user actions based on the current user context.
This comes across as nonsensical gobbledygook, but also a bit dystopian.
Seems like folks are reinventing the wheel, and trying to deduce what folks want to engage in with data and “AI”, rather than providing sufficient tools to allow the user to drive the narrative.
Pandora and Rdio and others solved the exploit problems years and years ago.
Doing music similarity with Echo Nest was great back when it was public. I did a project in grad school with it.
it's interesting how this continues a trend in across music/audio tech, such as hipsters insisting "i liked the earlier work better" or audiophiles obsessing over amps from the 1970s.
Artists, record labels, producers, technicians (mastering, vinyl cutting), distribution channels, etc. are all there just a click away.
Who did/does you favourite artist share a label with? Who else used the same recording studio? What other aliases does an artist use? What other bands/groups have they been in?
Combined with the embedded YouTube player on the release page it's a gold mine.
Every week I'll spend hours down the Discogs rabbit hole - often adding missing YouTube videos to contribute something for future visitors.
I've often wondered if Spotify could capture volume control data as part of a track to see if that produces better training data. But again its still too low fidelity.
I typically listen to full albums, but when I am working out or doing a road trip, having playlists for specific moods would be really nice.
- Playlists provided by spotify and tagged with various terms.
- User-made playlists (harder to find, but they exist, and many of them are fun!)
- Discover Weekly, Release Radar, Daily mixes - generated playlists for your account with 6-7 variations on how they are biased.
- Social features (see what your friends are listening to)
- Radios for songs/albums/artists
- Artist bios and "Fans also like" sections for each artist
- Smart shuffle on playlists (every 3rd song is a recommendation)
- A very permissive search box that lets you make mistakes (I hate competitors that punish me for writing things like "and" instead of "&" for an artist's name... I'm looking at you Tidal)
- Configurable search APIs to build your own funky queries or extensions
... and probably a few more tools to use when discovering music. An AI recommendation system will NEVER beat flexibility and giving users agency. I don't care that an AI takes into consideration that I went to the gym automatically. I care about having lots of options for discovery, each that is decent and biased towards a different style of exploration.
What if the user's current need is to not play music? To not consume yet more content? To not make them addicted to the content application?
How can we optimize for user wellbeing, and still make money? That's the question we should be pouring resources into
Sometimes I'll bounce off of a given artist several times, over years or decades, until the right track catches me when I'm "ready" for it, and then I'll enjoy discovering their whole catalog. At that point, the affinity is durable. I would love to find whoever is my next (e.g.) Steely Dan.
Then don't play music. You are asking for something that only works in a dystopia.
Either the machine tries to understand "what I need now" from partial information and will not play me music when I want to listen because it thinks it "knows what I need".
Or the machine actually is hooked up to my real time health data and possibly brain activity to actually know what I need.
I definitely don't want to have a personal computing machine thinking it knows what I want and deciding for me in such a way, or to have such access to my internal state.
There is a world where services are engaging, fulfilling and not harmfully addictive.
I love listening to one podcast episode each evening while I clean my kitchen. I do this each day. I would pay for this. I don't want this to hijack my brain and keep me up all night with content that requires an inordinate amount of willpower to put down.
There are plenty ways to provide content and make money off it without optimising for the best possible next item that keeps a person engaged until they fall asleep.