A reflection about automatic transcription of music (2018)
seventhstring.com
seventhstring.com
I would argue that there's a lot more people that can take the output of the MIDI and turn it into a reasonable score than there are people that can listen to the raw audio and turn it into a reasonable approximation of the notes.
Papers / dataset:
https://arxiv.org/pdf/1810.12247.pdf [New dataset and slightly improved network]
https://arxiv.org/pdf/1710.11153.pdf [Original Network]
With that background, I’ve come to believe automatic transcription is still so far from good that we’re better off creating tools that make it easier for humans to transcribe. That’s our philosophy behind Soundslice (https://www.soundslice.com/transcribe/), which combines a music notation editor with transcription tools. For anybody interested in transcribing music, I encourage you to give it a try.
If I were going to work on this, I would work on generating lead sheets from YouTube videos. Recognizing chords seems like a useful thing to solve?
It's also pretty easy to teach and learn for a practicing musician but much more difficult to teach a machine, since you run into issues with blind source separation.
Speech transcription is good enough provided we have enough preprocessing power, assume a single speaker, and know the language beforehand and have trained the model on a large number of previous speakers with the same dialect/accent.
In music, you don't know how many speakers there are (instruments playing), the dialect/accent (orchestration/chord voicing) changes on the fly, representations are non-unique and contextual, and artists intentionally subvert expected results to make good music.
Humans are just better at this and easier to train to do it than computers, for the moment.
Being able to separate individual notes of a musical piece into sharply defined buckets (keys of a piano) or one-dimensional subspaces (finger position on stringed instruments like guitars) simplifies the source separation problem a lot.
That representations are contextual and subject to interpretation by the artist is a harder problem (as discussed in TFA), but it should be possible to treat it separately from the pure chord recognition problem. (E.g. it would be easy to take notation and a matching MIDI file and then pretend that it's the output of the recognition step which the original notation should be recovered from.)
To use your example, finger position is not a one dimensional subspace on a guitar. There are between 1-6 ways to play a given pitch even assuming standard tuning on a six string guitar, which is not a safe assumption to make.
But the issue with chord transcription is one at the heart of source separation outside of VC demos, which is the causality problem that no one likes to talk about. To separate into N sources you need to know a priori that there are N sources to separate into. This is not a trivial thing to predict in chord voicings, where N changes and is not predictable. Then you need to make a best guess at which instrument the pitches fall into, which may be shared.
This is something that even humans fall victim too. Untrained listeners are very bad at quantifying sources in an ensemble, and even trained listeners struggle with notating chord voicings with decent accuracy.
Speech is dramatically easier. The reason is that language is meant to communicate and consequently carries a lot of redundant information that you can bring to bear.
Music often has no such redundancy. It may have themes that differ slightly each time, so nothing to lock onto.
> Recognizing chords seems like a useful thing to solve
Except that a single C note (especially on stringed instruments) may have many harmonics that also look like a C chord. The problem isn't straightforward.
It's not terribly difficult either. For plucked or struck strings, you can rely on the inharmonicity of the overtones to distinguish fundamentals from harmonics. Ensemble bowed strings are more tricky because the overtones are mode-locked, but we can rely on variance between individual players in both the time and frequency domain.
Polyphonic pitch detection is a bread-and-butter feature in many commercial music technology products and the industry is many years ahead of academia; I have absolutely no doubt that you'd be able to buy automatic transcription software right now if the market was big enough to justify the R&D spend.
I'm inclined to disagree with your assessment. Transcription is to polyphonic pitch detection what machine translation is to OCR. We are okay at the latter, which is what folks like Celemony make their living in. We still suck terribly at the former.
The last 20% will be fiendishly hard, as it depends on a mountain of cultural knowledge. A tool to help amateur transcribers to get ~70% of the way is probably within reach of current tech.
It's a shame we'll never get to hear Beethoven or Mozart playing their pieces themselves, in their intended locations and with their preferred instruments and seats and such. I suspect it'd sound much better than any other performer who's interpreted their work.
Here's a lovely representation: https://www.youtube.com/watch?v=G2tEVVeGCk0&feature=share
It would seem to be likely that given enough training data (labelled scores) you could make a net that took raw complex music in and generated that out.
I remember using Transcribe!, created by the author of this article, to try to figure out some of the fills in ZZ Top’s La Grange, which isn’t a complicated song exactly, but I can’t figure out how to make the fills feel right. Transcribe! Is great, btw.
Anyways, I’d be very curious to hear if there’s any existing work in this area. I feel like drum transcription could be easier in some ways than “pitched” instruments, but possibly harder in some unique ways too.
That said, I'm not sure yet if I'll be able to close the gap, and it's a hard problem to solve. And yes, like the author correctly identifies: for piano solo, or harpsichord.