Omnizart: Library for automatic music transcription
github.com
github.com
https://piano-scribe.glitch.me/
I built something similar which is a lot faster but on a large scale test the google software handily outperforms my own (92% accuracy versus 87% or so, and that is a huge difference because it translates into ~30% fewer errors).
Also interesting that the absolute timing of onsets worked better than relative timing - that also seems kinda bizarre to me, since, when I listen to music, it is never in absolute terms (e.g. "wow I just loved how this connects to the start of the 12th bar" vs "wow I loved that transition from what was playing 2 bars ago".
Another thing on relative timing.. when I listen to music, for me, very nuanced, gradual, and intentional deviations of tempo have significant sentimenal effects - which suggests to me that you need a 'covariant' description of how the tempo needs to change over time, so, not only do you need relative timing of events, you also need relative timing of the relative timing of events as well
Some examples:
- Jonny Greenwood's Phantom Thread II from the Phantom Thread soundtrack [0]
- the breakdown in Holy Other's amazing "Touch" [1], where the song basically grinds to a halt before releasing all the pent up emotional potential energy.
[0] https://www.youtube.com/watch?v=ztFmXwJDkBY, especially just before the violin starts at 1:04
[1] https://www.youtube.com/watch?v=OwyXSmTk9as, around 2:20
Rubato is everywhere in classical music, and understanding rubato is an essential part of any automatic transcription system that aims to show you notes in musically meaningful units of time.
[1] https://www.youtube.com/watch?v=h-eEZGun2PM [2] https://replicate.com/p/qr4lfzsqafc3rbprwmvg2cw5ve
I can't help but feel it is heavily impacted by ambience of the recording as well. The midi is of course a very rigid and literal interpretation of what the model is hearing as pitches over time, but of course it lacks the subtlety of realizing a pitch is sustaining because of an ambient effect, or that the attach is is actually a little bit before the beginning of the pitch, etc.
If it could be enhanced to consider such things, I bet you would get much cleaner, more machine-like midis, which are generally preferable.
For relatively "conventional" music, there are very strong signals of key like beginning and ending chords, and overall note distributions which will generally cluster around one particular scale. For less conventional music, this will be more ambiguous, but it would have been more ambiguous for a human listener too.
I found this link to be more helpful than the GitHub repo for understanding what it does:
https://music-and-culture-technology-lab.github.io/omnizart-...
The colab notebook is full of warnings and crashes with errors in the "Transcribe" box. Replicate.com does something but the results are garbage.
What am I doing wrong?
Which of the existing string music transcription libraries would fit the bill?
(I’m a good programmer in general, but still decidedly a layman in audio matters. Take my thoughts here with a grain of salt.)
Haven't tried it but maybe the following app could help?
> Trala uses signal processing, groundbreaking technology that analyzes the sound of your violin and gives you instant feedback on pitch and rhythm every time you practice. When you play the wrong notes, we’ll help you get back on track.
https://www.trala.com/ (I just remembered it being recommended somewhere.)
The algorithms as I understand them are grading you based on sustained pitch (adjusting for octave) and cadence. Generally they are easy to fool - you can basically say gibberish but as long as it’s within the expected parameters you’ll still get a good score.
They are still rather fun though if you enjoy that sort of thing.
All the examples given here, though, appear to be of a super-simple variety, with dead-simple chords, all notes with robotic, mathematically-simple timing over fixed tempos produced apparently by drum machines - like toy music, the kind of music that's no challenge at all to transcribe, and in the real world I wouldn't bother transcribing by hand as there's nothing to be learnt by doing so, as you can hear exactly what's going on without it. So, that's weird.
Transcription is a valued skill among musicians. The ability to hear a piece of music and write it down can be more than a little more difficult than transcribing speech.
Not sure if apocryphal, but Mozart is known an expert transcriber, able to hear an orchestral work and write down all the different instruments as they played their parts on just one listen.
In this case they mean to take music (audio) and write down in musical notation everything that’s being played.