Poor schemas, poor cataloguing: why music tagging sucks
sporks.space
sporks.space
Some good contemporary soloists: Alexandre Tharaud, Stephen Hough, Peter Hill, Jenny Lin, Isabelle Faust.
https://www.bbc.co.uk/programmes/b006tmtz/clips
Record Review:
https://www.bbc.co.uk/programmes/p02nrtty/episodes/player
More to explore here:
https://www.bbc.co.uk/programmes/articles/3cjHdZlXwL7W41XGB7...
I'm pretty sure I discovered Carlos Kleiber through Building a Library. Listening to his Beethoven Symphonies was like almost listening to different music, especially when comparing against say Karajan...
Elide the end, and you might get:
- Sonata No.[...].mp3
- Sonata No.[...].mp3
- Sonata No.[...].mp3
Elide the beginning, and you might get:
- [...] by Julia Fischer.mp3
- [...] by Julia Fischer.mp3
- [...] by Julia Fischer.mp3
Elide the middle, and:
- Anne-Soph[...]minor.mp3
- Anne-Soph[...]minor.mp3
- Anne-Soph[...]minor.mp3
I really hope that music apps would provide a choice for multi-lined, non-elided presentation of piece titles.
The elision itself is a problem; but so is the solution that involves horizontal scrolling of the information. It usually proceeds in _n_ character blocks and there’s a nearly zero chance the user will have the patience for the eventual completed reveal. Which movement of the sonata is this? I’m not waiting for the scrolling marquee. Table views are where the elision problem really surfaces.
From my own personal curation I just rely on memory essentially...
If you make the schema right, it's too complicated and people will not use it, because they just don't care that much.
If you simplify the schema, it's more likely people will use it. However, if it's too simple (artist/album/title/year), you end up with many many inconsistencies and duplication.
Finding something in between is nearly impossible.
With MusicBrainz, we've tried to design a strict schema that works for most things, but then you need to find people to actually enter data in that schema.
Wikipedia is on the other end of the spectrum, everything is free-form and some structure is slowly emerging from that, but it's far from universal.
Structured metadata is just hard for people to manage. Unless they are geeks and they really really care.
I feel like WikiData's support for music tagging is pretty robust and flexible, albeit hard to get people to enter data for
www.discogs.com
They have implemented a database with a pretty good representation of all the music releases (realizations, in other words), as submitted by volunteers (much like Wikipedia). How many of the questions posed here have already been resolved in a reasonable way by Discogs?
Then there was the late great Catraxx database, using much of the same relational structure as Discogs, but based on the MS Access engine. It can be used to write tags to audio files.
Personally I think the main draw of Discogs is the marketplace, even with the level of scalping it has.
I think a good 50% of my local music is missing from Musicbrainz.
When the release is present, almost always it's missing different issues for the same release (e.g. the white label test pressing release, the final release, the restamp from 20 years later). Discogs almost always has all of them.
Discogs includes lot of useful information: is that specific release single sided? Was it pressed on acetate rather than normal vinyl? What had it etched on the runout?
I use mp3tag connected to discogs (auth tokens)
It's incredible - i feed it files, it spits out glorious tags, folder names (perfectly formatted using regex)
The only thing that imo is annoying is getting catalogue IDs and using discogs + mp3tag to bulk/intelligently provide catalogue number.
Believe it or not, I had my entire process nailed when Bento existed. Then they got rid of it and I tried to cobble something together with Filemaker and it was always an awkward fit.
I'm not really aware if there's an open source or commercial offering out there for this sort of thing. I've come close a few times to just investing the hundred hours or so to roll my own web-based thing on a private server.
For example, I like jazz. I happen to enjoy listening to Toots Thielmann (a jazz harmonica player). There aren’t many well known jazz harmonica musicians, so a recommender system always gives me a bunch of other harmonica players from other genres. This is not at all a good recommendation, since the style of music and the style of playing in other genres is completely different (and not of any interest to me).
There needs to be a way of establishing “user input” as a way to weight the recommender algorithms better. Sort of like a search engine with “+” (more of this) and “-“ (less of that).
Then maybe recommender algorithms will learn and get better.
But I don't use music streaming services at all -- I run my own media server and stream from that. So what I started doing years ago is to ignore any existing metadata (even song and album titles) and enter it all myself. I've developed a system that meets my needs.
But it also means that every time I buy a new album, I'll be spending 15 minutes or so entering all the metadata for it. That sucks. I'd really prefer it if the online music databases got their act together and tagged everything at least in a consistent manner.
None of the fields in the databases can be trusted completely, but the worst offender of all is "Genre" and similar.
(They could implement it for most other formats, of course, but that would make their own formats look bad...)
Second, how to systematically tag the music based on the defined nomenclature? Without an automated system, it will be subjective and error prone. This problem could be maybe tackled on long term with machine learning.
It’s not cheap, and it takes some effort to fix some bad source data, but I’ve found it very rewarding and get a ton of enjoyment exploring my library now.
I added a bunch of songs to my favorites in Spotify and then hit the enhance button which is supposed to suggest other songs I might like. Many of the songs it suggested were duplicates of the ones I already had, but from different albums. It wasn’t able to recognize that this song was the exact same song in various albums of the same artist.
I can give you a very high level overview of why it's the case.
Simply put, they are different songs, and it's near impossible to recognize that they are the same song (at least not efficeintly across the entire catalog) due to a combination of any, or even all, of these factors:
- they are on different albums. The albums might not be attributed to the artist. Or be a compilation, and attributed to many other artists
- they may come from different sources and copyright holders
- the metadata may be wrong, or just different (metadata is supplied by copyright holders, and it's often ... weird)
- they may very well be considered different songs by most catalogs (a Japanese bootleg version that is 3 seconds longer is different from the European Best Of release etc.)
- everything is the same except some of the people on the record (e.g. arranged by a different person, so attribution and royalties come into play)
- ids, hashes, lengths, musical structure, or whatever internal systems use to identify, match, combine, and display music may all be different, or different enough to be classified as different songs
- there might be not enough music classified in the genres you listen to to present you with a large enhanced playlist. This is an issue for most non-western music because western music is largely understood, catalogued and matched for most of mmore-or-less popular genres. And even there it probably mostly applies to US nad British music. You're looking for African jazz? Good luck. Internally it's likely just one big lump of music dumped by music providers.
Only if you want your service provider to do the discovery for you.
I have AM and I don't need or want their discovery as I am quite proficient at doing my own discovery. This means that 3 of 5 icons on the bottom row of their ios interface are a total waste of incredibly valuable UI space for me.
Sure, you'd have to have labels like "Release Year 1999", "Track #2", etc., but this path can actually be very elegant and desirable at query time.
A few generic columns added to the label/tag table would allow workarounds to most of the edge cases. For example, if you added an "IsSequential" column, then labels like "Track #1", "Track #2" can be interpreted as such.
I think the dragon is trying to build a schema that directly represents the business of music. There are so many genres, cultures, varieties, etc., that you would go insane before you had everything properly covered.
See Every Noise: https://everynoise.com/
Is modern chamber music different from modern classical, and why? How different are Polish free jass and free imporvisation? Canadian black metal and Norwegian black metal and Dutch black metal?
To some, there's no difference. To others, there's a world of difference.
Here's more on the difficulties: https://everynoise.com/EverynoiseIntro.pdf
But the bigger issue is that while it’s pretty good at finding songs that sound similar it’s not great at finding songs that are musically similar.
But nonetheless I really appreciate a novel recommendation algorithm that’s not based on popularity. I’ve gotten good recs from this site with with less than 10 monthly listeners which is super cool — I’ve never been so underground.
Then it was improved to make the matching even more accurate, and it was -- but that made it less useful to me. It turned out that the inaccuracy was a feature, not a bug. When the matching got too good, I started getting music that was too much like the stuff I told it I liked, and I stopped finding so much great stuff that would have escaped me before.
1. what resources would you recommend?
2. Is there a "gold standard"?
Article already mentioned some great terms to look up and explore.
One example was Rammstein's Untitled album, it literally doesn't have a title. How do you tag that?
(if anyone has the link to the article, I didn't add it to my Wallabag collection...)
Thanks!
Track order # is part of the metadata, so it shouldn't have been an issue. But the music app I was using on Android at the time would crash whenever it got to a letter which was repeated (so the 2nd R E or T).
If you look up the album on streaming services, they use the naming format "TRACK 1 R, TRACK 2 E" and so on.