Thank You for 20 Years of Discogs
blog.discogs.com
blog.discogs.com
As a vinyl collector, thank you Discogs!
https://github.com/beetbox/beets
it takes a while to take a large library, but that's mostly because of intractable problems with ambiguity in metadata searching. I think it's about as fast as is reasonably possible.
Every new release that gets added to my collection will spend a few seconds in Meta so that I can check and edit the meta data and rename the files in a consistent way. Obviously you can also edit multiple releases at once.
I used quite a few tagging apps over the years, Meta is without a doubt the best of them.
- Discogs seems to have the most complete release data.
- MusicBrainz has the best organized encyclopedic data. For example they have entities for specific instruments, relationships between artists, entries for supporting roles like recording engineer, etc.
- Wikipedia still has the best common-sense/canonical info about singles. For example, Wikipedia will tell you in the first sentence (or infobox) "song X was first released on album Y in year Z." Answering that can actually be quite challenging on other services.
- Spotify is the go-to for finding useful qualities of individual songs (due to their acquisition of Echo Nest). Things like popularity, danceability, energy, etc.
A long-term project of mine has been wrapping all these up in a more approachable API. So far that has only resulted in this big GraphQL schema project: https://github.com/exogen/graphbrainz – but the idea is to build on that and make it even simpler.
https://developer.spotify.com/documentation/web-api/referenc...
I used to do a lot of data mining on data from various musical sources. Mostly with the goal to make the discovery process faster.
Sad fact: For a while Discogs had an experimental feature aptly called "Tracks". It allowed individual tracks, not just the whole release, to be viewed as separate entities along with references to all releases on which said track appeared. A database of canonical tracks.
It enabled me to see if that great track from an ultra-rare and hard to find album perhaps appeared on a compilation or any other much easier to buy release. Sure, you can achieve similar results by simply searching, but it's much more cumbersome and downright difficult when it comes to rather generic track titles leading to countless irrelevant results.
Sadly the feature was abandoned: https://www.discogs.com/track
Can you explain what that refers to? Because if I search for that I obviously get WW2 related topics instead
A ‘45’ fits only a song each side, so usually for releasing singles.
And ‘German’ means German.
My family who grew up with vinyl records are obsessed with figuring out when knowledge of that technology will really start fading.
But they also have a pretty good API with generous limits. Really need to find the effort to restart work on my "spotify shuffle/playlist" style app that uses your collection to build out a play session. Had a lot of momentum at the start of covid and just fell off.
Discogs not only has an API [0] but they also provide regular database dumps for free [1]. Last year, I converted tjeir database into SQLite and queried it in various ways to discover new artists or releases that I might be interested in. Ultimately, I realized how much of my favorite music was put out by labels that were headquartered in UK flats that have since been re-leased.
[0] https://www.discogs.com/developers/ [1] https://data.discogs.com/
Also curious what labels are on that list of favourites!
For example, lots of popular artists in Japan (not trying to cherry pick something obscure; Japan after all has the second largest music industry in the world) don't even have an article there, let alone a complete discography.
This is obviously due to lack of volunteers, or their volunteers have limited areas of interest, which is totally understandable. But I think some basic automatic systems to at least add the new releases from major labels is needed for such database.
I personally have a better experience on https://musicbrainz.org/. It suffers the similar problem mentioned above, but much better. At least most of "mainstream" artists I checked, it has complete or close to complete discography. It is also a more "ambitious" database, which has data for each track, recording (my fav part), etc.
I guess we can say that all these crowdsourced music databases have obvious priority or bias, despite they are appeared to, or want to, be a comprehensive one, which isn't very realistic just due to the sheer amount of the music produced. In Discogs' case, even though it gradually expanded to "everything", its strength is still at its core (where it started): electronic, hip-hop, rock, jazz, (in this order, named in the article), but probably not pop.
One probably would have much better luck looking for something in one than the other, totally depending on the genre. Like, if I'm looking for game music, I definitely will check VGMDB first.
Of course, it's not limited to music either. It can happen in much, much smaller fields. One particular case I want to mention is that there used to be three wikis for a single game, Dota 2, and each have their own strengths (there are still two today!).
I think Discogs mainly caters to record collectors (vinyl etc.) which is not often covered by other sources. While honestly musicbrainz info is more common/shared.
This makes sense, I do have better experience when looking for vinyls, regardless of country of origin.
Still think they should at least have a basic dataset for newer CDs, though (to be fair, MusicBrainz's new music catalog is often lagged by a few months too.)
They don't exist. Look at SoundExchange, the US organization that handles music royalties from streaming. They don't even have complete records on the releases for which they collect money! You're talking about a myriad of Hollywood-accounting companies (if that) that were fired up and dismantled willy-nilly, for decades.
Oh, that Japanese label that put out two 45s in 1962...where would that be recorded except in the collections of music fans who own them? That's what Discogs is based on. You want those records to be in the database? Buy them and put them in. The fact that you don't is what you're describing as "limited areas of interest."
I'm saying so because I know plenty of 3rd party fandom/crowdsourced databases are doing exactly this, for books, CDs, etc. The data obviously need some manual checks later, but this step saves lots of time.
Again, Discogs already have a very complete database for old releases, especially for vinyls. But what I'm talking is that they are lacking for newer ones (CDs or digital releases) and just provide a systematic (and tried) way to improve in that regard (considering their goal now is "everything").
For something in 1962? Sure you're right. But you don't really need some fans to own a copy to prove the existence or get information for a hit album in 2010s.
So all the individual credits for e.g. the song "The Boys Of Summer" by Don Henley could be updated in one place, and then updated on all the releases it appears on.
That would make the site extremely useful to see who played what on which songs, who produced and mixed it etc.
You CAN find that information for songs now, but since a track might appear on hundreds of releases, it's impossible to know easily which release has the credits for it.
Information propagation is good IF AND ONLY IF the information was correct in the first place. But it's a big problem how much incorrect information already exists in music databases (including copyright databases, which make a living off providing supposedly correct information), and blindly copying information between sources propagates such errors.
Preserving the diversity of sources, and only judiciously updating information, at least gives a chance of evolving toward a more correct representation of the world.
Discogs, for instance, seems fairly accurate for release dates (Something that e.g. AllMusic can't be trusted for at all). And discogs is a wonderful source for photos of album front and back covers and record labels, which by definition are first hand evidence (even if they themselves not infrequently contain errors).
Information on e.g. track composers is unreliable, and musician and label information is sometimes a mess (multiple entries per musician, inconsistent attribution of releases to label / country sublabels, musician / band). All of this is hard to avoid in a huge database like this.
I have not worked with Musicbrainz as a source, but it does not have a reputation as being particularly reliable.
This seems like a roundabout way of saying it has a reputation for being unreliable?
Do you mind clarifying? I know there's certainly issues of completeness, as with any user-compiled library of data, but I haven't heard charges of unreliability levied against it before.
But let me try: An ideal data source has a traceable provenience to a reputable authority. E.g.
https://iswcnet.cisac.org is in the business of establishing songwriting royalties.
The Discography of American Historical Recordings https://adp.library.ucsb.edu/index.php/basic/search has very sound scholarship behind it and documents their sources (usually ledgers kept by the recording studios).
The Encylopedic Discography of Cuban Music https://latinpop.fiu.edu/discography.html seems to be compiled mostly by one scholar, but he seems to have an incredibly detailed knowledge of the field.
Musicbrainz, in contrast, is a crowdsourced collection that is a mile wide and an inch deep. Within literally seconds I found a bogus artist entry: https://musicbrainz.org/artist/729f9e2f-c993-40b6-ae20-bd21b...
I'm willing to bet that the the last album shown for James Carter is not by him: https://musicbrainz.org/artist/a8880ecc-10d8-4492-b17e-02715...
This one has a typo in 13, and 12 is completely wrong: https://musicbrainz.org/release/81c40619-26b9-421a-906a-4ef0...
And all of this is just a few random checks. To be sure, the sheer volume is not without value to start a search (Quantity has a quality of its own, as they say), but it should probably not be taken as a sole source of authority without a backup source, and one would have to be careful that the backup source did not, in turn, take musicbrainz as its source.
Your argument seems to verge on absolutism, as in "if it's not perfect, it is worthless", which makes me feel like it's a little dismissive of the value that musicbrainz has managed to build.
I'd like to argue that most if not all data collections like this, including Discogs, Musicbrainz, web stores like iTunes or Google Music or Amazon are always going to have some data quality issues. It's the nature of the beast. I've seen data quality errors even in Apple Music and Spotify, where you'd think they have vetted data from labels directly.
MB has not one scholar but thousands of enthusiasts contributing to it. Some are very methodical and provide source links with each edit (creating, as you say, traceable provenance to a reputable authority). Some are newcomers and just throw data in, and need to be educated on how to do it better.
The great thing about MB is the powerful data model, and the UI that makes it possible both to fix the invalid data, and to trace the history of the edits, with links to the sources used for the changes. The other great thing is that it has, so far at least, somehow managed to elude most of the attention of vandals and trolls that plague places like Wikipedia. It feels like most if not all contributors to the site act in good faith to their best knowledge (lacking though their research skills may sometimes be), and when their submissions are lacking, they are easily improved.
Finally, I do agree with your ultimate statement - none of these data sources should be taken as the sole source of authority.
To a very specific group of people collecting old electronic and dance music, it is indeed invaluable and unique.
edit actually the database is available at data.discogs.com
A friend of mine even found an old LP of his band there after he lost the last copy he had. (They disbanded decades ago and only pressed a small amount of copies)
Sometimes the disc isn't recognized or you need cover art and discogs is there for free.
So no, I'm closing this window and not even reading the content.
I thought the EU court clarified that this is not what they meant about reasonable opt-outs being required.
In the beginning it was very much a community driven project and it seemed like the Discogs database would remain open like e.g. Wikipedia with the GFDL, so after pouring in a lot of work it was very disappointing when the project was commercialized and the database closed with a restrictive license.
I really wish they would include the marketplace statistics, because from that one could estimate popularity of releases. But since that data is not crowdsourced, it's fair enough that they keep it private. Of course, one could scrape it from the web, but that is impractical and unethical to do for all the millions of releases.
The database is the total accumulation of literal decades of work with almost nothing thrown away, so it's not the cleanest to navigate for a newbie, but it is deep.
Thank you, Discogs!
An excellent feature is their wishlist. Add releases you would like to obtain, and you can get an email "newsletter" with new/good deals on these releases. You also get to see if one seller has multiple wishlist items to help reduce shipping. Or just browse through the seller's other listings and add a few cheap records without raising the shipping cost.
I wish Metal Archives weren't offensively lazy sacks of shit about their own data and made an API. I'd work on that shit for free for them.
Some interesting musicological statistics can be computed using their public dumps. For example: http://hdl.handle.net/10230/32931