Using Spotify to measure the popularity of older music
poly-graph.co
poly-graph.co
A popular cover can probably drive a lot of popularity for an older song on Spotify.
One nitpick: "Iris" by Goo Goo Dolls was ineligible for the Billboard Hot 100 for a long time due to their rules that songs had to be sold as a "single" to be counted on that chart. It was #1 in terms of radio airplay for a huge chunk of 1998, so it's not really accurate to say it didn't chart highly at the time.
This doesn't add a single thing to the discussion, but it's a great cover too.
Regardless of the age of the listener, I expect newly released songs to have higher playcounts (in fact, I plotted this curve, but it was too high-brow for the Internet and the audience for which I was writing).
That said, if I managed to cut the data by age-bucket, I do think that the results would shift toward the music with which you grew up.
Few slight bugs on the Present-day Popularity of Five Decades of Music, Dream On appears twice in the 70s with the same listen count(73 & 76). Also Blink-182, 1999 is showing in the 00s. All I Want For Christmas Is You — Mariah Carey, 2000 is showing in the 90s.
Why lasting popularity as a measure of timelessness?
How do you account for longer trends? Some of Bach's children were more popular than he was for quite a while.
We only have two data points in this work: today and release date. So longer trends like the one you pointed out might be lost in time.
But, very interesting data nonetheless! I'm loving it!
Is it possible to filter the data by listener age? I wonder, because a lot of these songs are in play lists of mine from the 80s / 90s (I grew up in the 80s / 90s). Maybe spotify's user base is older than suspected?
Also, it would be interesting to plot billboard rankings vs spotify rankings. Possible?
The last point, Billboard vs. Spotify, is in the second to last chart. Check it out :)
I used D3 to create the charts, as well as some additional frameworks (Jquery, Waypoints).
Also so happy to see that "You Got Me" by The Roots got 6 million streams in 2014. That is the definition of a future timeless track.
Did you manually retrieve the play count for each track or is there an automated way of doing it?
I'm sure Nirvana has many #1 hits, but why did Smell like teen spirit become the poster song for Nirvana?
The same can be said about Oasis, who at the time was insanely big and held several spots on top 10 lists for months. But maybe they were bigger in Europe than the U.S. And maybe European fans are driving the spotify listens?
Or is it just plain and simple data gathered from Spotify?
I'm not saying that there is anything wrong/bad about the results. But without knowing the details on how the data is collected, it's hard to read anything from the results.
No Diggity is a great song, but the song it samples might be even better: Grandma's Hands by Bill Withers.
Second of all I do dislike texts on data which lack information on where the data comes from.
I can think of ways to mine present day play counts from Spotify (while not working there) but I wonder where did he get the daily counts from he used in the last chart. Any ideas?
Furthermore I doubt that Spotify is necessarily a good indicator on how songs are being perceived in the long run. Especially b/c there are local platform-specific attractor dynamics at play.
The data is pretty clear in terms of source...Spotify in 2014...Billboard data via Whitburn.
The data was directly from one of Spotify's data partners.
Yea Spotify isn't a perfect indicator. This is the best proxy for present-day popularity that I can think of. I could have create an index that abstracted several data sources, but that would have killed the readability of the article.
Just switch it on ... it's your site, isn't it?
> The data is pretty clear in terms of source...Spotify in 2014
That's not the "source" that's just a value of the time dimension.
> one of Spotify's data partners
well, you could have given that information in the text - if you talk about data, you gotta talk about where you got the data from.
Nonetheless the statement is still pretty obscure. Who is that "partner" - is it a secret?
Why don't you just dump the data on GitHub?
> I could have create an index that abstracted several data sources, but that would have killed the readability of the article.
I'm not sure if that is the true reason why you chose not to do it - but if so, then it is necessary to be transparent with assumptions, abstractions and simplifications, right?
I know that this undermines the credibility of the article, but I'm optimized for readability and storytelling, not to build a full-proof argument for timelessness. There's a million rabbit-holes that I could have gone down to make a much more solid case, but I decided to present the data and let the reader draw conclusions (kinda like I did with the hip hop/vocab piece: http://poly-graph.com/vocabulary.html).
I also realize that one could argue that this is a terrible way to approach a writing/data-analysis project. Assumptions and simplifications are important to highlight. But I weighed the options and decided to focus on accessibility.
Happy to discuss the pros/cons of this further :)
In which case, I should voice my opposition to the suggestion that we non-USians (eg. I'm an Australian in China) should communicate (even in our own language) with hat tipped to US popular culture because (inertia of Colosseum-fawning masses).
Here's a contrary view: I believe that intelligent people tend to respect and encourage diversity because it's both more interesting for them ("are we nought but latter-day curios for the coming AI overlord?") and because many fields of science (chiefly biology) show us strength in heterogeneity. The parent's comment was, I believe, offered in this spirit.
The lyrics are a little fluff, but Ironic has a great chorus and catchy hooks. It's a classic pop song.