How Spotify ran a large Google Dataflow job for Wrapped 2019
labs.spotify.com
labs.spotify.com
[0] https://community.spotify.com/t5/Live-Ideas/Account-Change-U...
One thing that I wonder about is how much work could they do to collect this data on a forward moving basis. Often I see huge lookback jobs that answer predictable/static questions -- prime candidates for aggregation during ingest.
Seems as though users were pinned to some general playlist that had characteristics similar to listening habits? Still hats off from an engineering perspective. I as well wish there was more technical detail provided.
The year recap playlists though are fun personal snapshot of time.
Meanwhile in Denmark or Poland there is very little in terms of data limits.
Eventually, Spotify released its official mobile apps and a web player so the project had no use. But it was fun times, it was really marvelous how anyone could find their favorite music from the service and listen to them in good quality without a torrent connection.
Nowadays, I think all those friends who used the hack are Premium subscribers.
Overall the suggestions are good when they’re actually derived from what you listen to, but stuff like this really bothers me. Last night I saw some of it creeping into the discover lists which makes me wonder if the good recommendations are coming to an end. There’s certainly money in it for them in the short term.
I completely agree.
> In this case there can’t possibly be people arguing for their own datacenter over cloud.
Devil's advocate time: This solution was great for the cloud because it was designed for the cloud. There might be equally good or even superior solutions designed for on-prem or even on-device computing. For example, this ceases to be a big-data problem if you are simply aggregating listening metrics for a single user on a single device.
If Spotify leveraged my phone to calculate these statistics of my listening history (owned and stored locally), this article would have been written about an app update.
No need for a massive ad-hoc job with high-bandwidth round trips, just a simple app update.
It’s funny to imagine how engineers of the future might look back on our pride over this kind of computing similar to how we look back in horror on how wasteful we once were with mining oil back in the 1910s, etc.
Then the article would be about the challenges of battery life on users' phones, and trying to coordinate listening history on PC vs. phone.
The article on coordinating and compressing listening history (the particular challenges of distributed schema evolution at the “edge”), would have been a much more interesting article to read, IMO.
Also, I know you probably weren’t very serious about it, but I don’t think that a few SQL queries against “thousands of data points” (temporal rows, reading between the lines) would be a significant battery life drain! It would have still been interesting to see that benchmarked. But “big data” is cooler, I guess. :)
Of course this all depends on the level of detail they want to store, it could be a uuid, a tstzrange, and some Booleans about whether the song was liked, downloaded, etc.
Every year (or once you reach some storage threshold) you could “compress” this information by aggregating rows by song, and throwing away precision on the time stamps, until you’re just left with a uuid, full/partial play counters, and dates that the song was liked/unliked, downloaded/removed, etc. You could give users the option to modulate the level of detail in the records, to trade off storage constraints against recommendation UX.
It’s a set of constraints that differs greatly from a huge ETL job, but my point is that this kind of edge work leads in interesting directions, too :)
Definitely. Given that they're doing this every year, it seems perfectly plausible to do most of the work in an incremental or streaming fashion.
We've seen that in every industry including healthcare. Every health crisis now takes us back to field hospitals.
Space-wise, yes, but users are likely using multiple devices and may have switched phones, reinstalled the app, wiped data etc.
Then you have to consider that the scripts would have to be individually written for each platform, and would have to be careful about power consumption, CPU usage etc., especially on mobile devices. And there's not just data mining but also video encoding (for the stories).
And then there's this part:
> To bring you a Decade Wrapped, we had to process these data stories over 10 years’ worth of data for all of our monthly active users
I was under the impression that the stories were live graphics. They certainly where on PC, as I had issues running the WebGL because of my script blockers.
Another idea is to run probabilistic queries instead of exact ones, could bring down costs way more.
https://labs.spotify.com/2019/11/12/spotifys-event-delivery-...
Not sure if a better title is warranted ("How Spotify ran its massive Google Dataflow job for Wrapped 2019", "How Spotify ran one of the largest Google Dataflow jobs ever for Wrapped 2019"?).
I always tell startups not to use superlatives on HN. Modest language sounds stronger.
edit: Did a search, seems like there's quite a few problems (only playing recently added songs, only playing 100 songs out of the playlist, etc.). I know google music has also had long standing issues with shuffle play - and in fact I left it over these kind of issues. Is it really difficult to implement a shuffle?!
This may, of course, have changed. My experiments while (badly) implementing librespot's shuffle functionality were a few years ago now.
When I use the Amazon app under the same conditions, I often hear a track I haven't heard for a long time. Which is what I'd expect when random sampling from 200 hours of music.
(I don't use playlists, as they're simply too much work.)
Simply put when you shuffle from all of you liked songs you will mostly get the same tracks over and over - some tracks will stay hidden forever, - pretty weird and annoying.
It seems to stem from issues in relation to this post, ie. sql queries and caching to prevent too much CPU use on their end.
They perform some analysis to increase the "perceived randomness" - e.g., if the true random seed picks the same artist twice in a row (totally possible), pick another song by a different artist, or else people will perceive the shuffle as not "random" enough.
Unfortunately I don't have the source for this right now, but I'm sure someone will hop in and provide it if I'm wrong about this :)
That was hugely frustrating, but we would get user reports of the random button being buggy when e.g. the user gets 2 tracks of the same album/artist one after the other.
Of course that can happen if we truly randomize your content !
So we switched to a pseudo random algorithm that tries to have consecutive tracks from different album/artists.