Spotify Optimized the Largest Dataflow Job Ever for Wrapped 2020
engineering.atspotify.com
engineering.atspotify.com
> The intuition is that for datasets commonly and frequently joined on a known key, e.g., user events with user metadata on a user ID, we can write them in bucket files with records bucketed and sorted by that key. By knowing which files contain a subset of keys and in what order, shuffle becomes a matter of merge-sorting values from matching bucket files, completely eliminating costly disk and network I/O of moving key–value pairs around.
I'm actually surprised that this should be regarded as "novel" in data science.
It reminds me of something in Eric Raymonds "The Art of Unix Programming" (I don't have time to find the link right now) where it discussed an approach from the earlier days of Linux filesystems where you had a limit on the number of iNodes that could exist in a single directory and corresponding performance. The work around was to create a subdirectory structure to store files based on the filename. But then you tended to get many files starting with the same characters all in the same directories. What turned out to be a better way to distribute the files evenly in the directory structure was to take the first and _last_ character of the file name and use those to create the subdirectories. This way you were more likely to spread the files evenly across the structure.
Interesting. I have been pondering over filesystem performance and inode limits in servers/home-servers since a long time. This seems useful infomration
Index organized tables in Oracle, clustered tables in Mssql. "Intuition" in modern big data world :)
Before worker nodes had as much memory as they have now, almost everything needed to use small buffers and spill to disk. BDB (Berkeley DB) was an extremely common tool for doing out of core data operations. Because the ETL tools I was writing needed to run on machines with 512MB of ram, it required out of core algorithms. We easily had jobs processing 10-20GB with only 512M of ram.
I am sure I am missing something, reading the paper now.
http://kth.diva-portal.org/smash/get/diva2:1334587/FULLTEXT0...
[1] https://news.ycombinator.com/item?id=25839399 [2] https://news.ycombinator.com/item?id=25064636 [3] https://news.ycombinator.com/item?id=24699908
I used to think that Reddit was bad in this regard but to be honest it mostly affects the big subreddits, the niche and small ones still have a high quality community. HN became pretty much like the biggest subs on Reddit.
This comment thread is in its own category of low quality discussion.
Negativity bias prevents you from seeing that 95% of the homepage right now is technical/nerdy with a lot of high quality corresponding discussion.
When political/social issues hit the homepage, they often slide off quickly if the corresponding discussion is of low quality (has many downvoted comments).
HN is certainly not perfect but just focusing on the parts you don't like prevents you from seeing the bigger picture.
Mod. There is a question of how much one moderator can do against the tide. HN really needs a couple of full time paid moderators, with their salaries covered by the zillion dollar YC bank account.
This definitely drives people to comment on other things.
My gut feeling screams they made a problem themselves in the first place which they then "solved". Similar to a "solution running around looking for a problem" type of deal.
From what i undestand, Spark has the same feature built in. If the planner knows that the source data is partitioned and/or sorted appropriately, it can skip shuffling/sorting it, instead having each executor directly requesting the one file it needs.
It's a nice optimization, but it's not game changing. You often end up having to shuffle anyway, as you are joining on a different key, or for performance reason you need more executors than the set amount of partitions, or the shuffle needed to write the data doesn't justify the savings on the readers.
Maybe it's better with their additional optimizations? Spark does not do those, mostly.
I agree with you that it would be interesting to know, I just don't think it's realistic for them to release that information.
> "We estimate around a 50% decrease in Dataflow costs this year compared to previous years’ Bigtable-based approach. Additionally, we avoided scaling the Bigtable cluster up two to three times its normal capacity (up to around 1,500 nodes at peak"
The official Spotify Engineering Tweet similarly only makes mention that this is Spotify's largest dataflow job ever: https://twitter.com/SpotifyEng/status/1359887825047613442.
I'm fairly sure a similar accidental unsourced exaggeration was made last year.
Maybe the title should be Spotify Optimized Their Largest Dataflow Job Ever For Wrapped 2020?
If I am reading things correctly with Kafka the workflow equivalent to what's written in the article would be to have your producer produce via hash-based-round-robin (the default partitioning algorithm) based on the key you are interested in into some topic and then your consumer would just read it and your data would already be sorted for the given keys (because within a partition Kafka has sorting guarantees) and also be co-partitioned correctly if you need to read some other topic in with the same number of partitions and the same logical keys produced via the same algorithm. No?
> our data would already be sorted for the given keys (because within a partition Kafka has sorting guarantees)
It's been a while since I used Kafka but I don't remember "sorting guarantees". Consumers see events "in order" based on when they were produced, because each partition is a queue.
I'm learning proper data flow in real time as I look to transition ETL of product data into Postgres to a more applicable system.
Finding the right learning resources is difficult! Cheers.
I enjoy it anyways, and Spotify is still a great service for now - I wonder if it'll meet the same fate as Netflix at some point, with publishing houses going for their own streaming services instead.
I don't think so because the way that people consume music and the way they consume films and television are very different. With a film you might block out a few hours to watch that specific provider. With music you're more likely to want to interleave content from several providers at the same time (eg a playlist). Unless all the providers are available on the same platform it wouldn't work well.
There's also a common use case where people will just play a specific artist for an hour. Or even an album.
Frankly, I hope services like Spotify don't disappear. It's a great loss to consumers just how fragmented video streaming services have become. I'd hate to see the same happen to music as well.
My wife uses them to shuffle singles by specific artists (she’s more into pop music).
I’d wager if my wife and I both coincidentally follow the same pattern despite doing so for different reasons, that it’s then likely a more common pattern than first assumed. Please also bare in mind that I’m not suggesting our use cases are how the majority of people consume music, but I’d be surprised if it was small enough to be a rounding error.
And then I worry skipping it is going to train the algorithm that I don't like Pink Floyd.
They probably meant listening in random mode, and Spotify randomly choosing a track that doesn’t make much sense outside the context of its album.
It's not particularly actionable to know you listen to more music on a Saturday than a Sunday, but it is mildly interesting for those who are curious about such things.
Spotify itself is not a data product.
The first rule of avoiding disappointment is managing your own expectations to align with reality.
That said, you can export more granular data Spotify has on you from your settings page[0] if you want to do your own deeper analysis on your trends and usage, etc..
And merge joins from sorted data? Joins have been done that way since the punched card days on mainframes (and by any scaled data system)
From literally the first sentence:
> from our largest Dataflow job
> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that."
https://news.ycombinator.com/newsguidelines.html
The hackernews title and the article title say "the". Critizing this clickbait is more than warranted.
But I think it's a reach to call this "clickbait".
Article titles are shortened all the time, and you can't expect them to have all the context in the title.
However, one should reasonably expect people participating in a discussion about an article (particularly when posting criticism) to have actually read it.
Complaining about something that is provably false in the first sentence of the article is the bigger sin here, is it not?
I've run into artist naming collisions enough to know it's a thing but I've never seen anyone (ab)use it intentionally..
Neat, but still annoying I'm sure. =)
I wonder if some of the data in the "We worked with the maintainer of these data sets to convert a year’s worth of data to SMB format." step got corrupted or just wrongly converted/lost.
I'm not sure how else explain that I have to google artists in my top 10 because I never heard of them.
My email is me [at] param.codes.
What’s your point?
Other than that, I'm not surprised by what they log. Virtually every company stores search queries, oauth grants, play history, ad interactions etc. Doesn't make it right of course.
They're complying with GDPR. Isn't that a good thing?
The "scary" thing in that tweet is that they store the manufacturer of their bluetooth headphones?
It connects to your Spotify account and generates a nice public page with your stats (top artists, top tracks), playlists and etc.
You can reserve your username now: https://volt.fm
* Not overwrite/delete my listening history everytime I switch devices
* Allow tabs, or some way to resume what I've been listening to in different contexts
* Option to open only one instance, instead of having multiple instances that mess with each other
* Playing local files crashes/not working on Linux
* Change playback speed, not just for podcasts
* Jump back/forward, not just for podcasts
* Have some visibility when the song was last played / play count
* Liked songs not always appearing in search results
* Sorting search results not working
* Add basic functionality to the dbus interface (e.g. seeking)
* Ability to report songs (e.g. wrong titles/badly split tracks/etc.)
Nevertheless, they decided not to spend a dime to work on the Linux client ("Spotify for Linux is a labor of love") [1]. The Linux app is an electron app, so there's relatively little effort required to maintain it.
Instead, they paid 100 million USD for Joe Rogan, and added video features that nobody wanted and asked for.
Take a look at /r/JoeRogan [2], and see just the extent of how much reputation damage it caused.
Take a look at /r/Spotify [3], most of the complaints are about basic functionality not working.
MPRIS [4] is a dbus API standard for controlling music players. End-users don't know or care about APIs. End-users care about having music widgets, but you can't do that without it. End-users care about keyboard accessiblity when they need it, but you can't do it without those APIs. And Linux users specifically care about programmability.
Limiting the API means that not only do I have to deal with the buggy app, it's prevents anyone else from building something better.
It's been 10 years since the Linux client came out. One engineer could fix these issues in a couple of months, but it's just not their priority.
The only saving grace for Spotify is that all their competitors are worse.
1. https://www.spotify.com/us/download/linux/
2. https://www.reddit.com/r/JoeRogan/
To this day when I want a recommended playlist based on my taste/history, I always use last.fm because it's just plain better. Why? The "Discover" etc playlists on Spotify are just crap.
I wonder if adding 'as much music as possible' would drive any growth without the AI music discovery stuff. There has to be a diminishing return to adding new songs - if Spotify adds an artist that only a few thousand people have heard of then that's only going attract a few thousand new customers at most. No one else is going to listen to that artist unless Spotify recommends the songs to people who might like them.
The Spotify end of year wrap is amazing marketing for them that people absolutely love and share widely.
I think Spotify's "moat" is largely the analysis they've done on all the listening data they have, and their ability to provide that to the music industry as a product.
They share rudimentary summaries with customers via stuff like Wrapped but I have to imagine they have much more detailed and robust data products for the industry...
If they're just a giant Dropbox for MP3s with a music player sitting on top, they don't really stand a chance...
I miss being able to do something simple like listen to music or watch a movie without all my actions being recorded and saved. So I'm back to buying physical media and DRM free downloads.
I'm convinced that it is now important to hold on to older appliances that work without internet access or data collection this plus right to repair gives me hope for the future.
If that’s the case, why not just sign up using a single-use email address, pay via gift cards purchased with cash, and if you’re really concerned, use a VPN?
I don’t think it’s a surprise to anyone that Spotify collects telemetry data on its users.
The main reason myself (and many others) use Spotify is because they use this data to recommend new music that I will like.
It's not meant as slander. It's just to show others 'here is something that's true that you might not have thought of or have thought of but haven't thought of it in all it's 250 MB glory. Decide for yourself'.
Folks who wanna try something different might like 'Radio Paradise'. It's human curated music. Have wide platforms support [1] but also regular stream URLs [2]
[1] https://radioparadise.com/listen/options [2] https://radioparadise.com/listen/stream-links
Why should I be worried about a company knowing what music I listen to?
Or run modern, up to date FOSS equivalents on machines you control.
I’ve migrated more and more services like that and I’m slowly but surely building my own “cloud”.
Then it's not always good, right or even close sometimes, but it's not like it's a hidden feature.
Although it is strange to leave Spotify in 2020 citing concerns about Spotify using your data, as it has been going on from at least 2015, possibly earlier.
2020 was the year I really started to question why I was taking such care with my data in some ways but not others. The Wrapped 2020 was a bright reminder to me that if I want to take my privacy seriously I need to look at everything in my life that collects data. Simple as that.
I don't think spotify is wrong or evil to collect the data. I actually think it is a great product. I have just decided that I want to leave as little data around about my daily activities as possible.
As it happens I don't use the Discovery features very much and it turns out there are still some enjoyable FM stations where I live. When I want background music I turn on the radio. When I want something more specific I play an album or playlist from my collection.