Pitch Detection with Convolutional Networks
0xfe.blogspot.com
0xfe.blogspot.com
Here is the video archive of a recent conference on the topic.
https://ismir2019.ewi.tudelft.nl/
There are some FOSS applications e.g. https://www.sonicvisualiser.org/ but I'm surprised how bad the results of the analysis are. Intuitively, it seems such a simple problem.
https://manual.audacityteam.org/man/audio_track_dropdown_men...
Side note: cepstral processing is going to be a lot more effective than spectrograms, and preprocessing is cheaper than training the ANN.
Not sure why you think so. Almost everything you suggested (missing fundamentals, noise, vibrato, reverb, velocity, distortion, etc.) can be synthesized with tools like sox. It worked very well for me. :-)
> cepstral processing is going to be a lot more effective than spectrograms, and preprocessing is cheaper than training the ANN.
I did do this in my initial attempts, and found no improvement over spectrograms. Turns out NNs can learn log nonlinearities quite easily. (EDIT: to be precise, I calculated the mel-cepstrum and fed it to the network.)
The mel scale cepstrum is inappropriate for pitch detection. You want to use the cepstrum without scaling frequencies. Take the fourier transform, normalize, take logarithms[+], then (inverse) fourier transform in the frequency domain.
The advantage of using the cepstrum for pitch detection is that most signals you're looking at will be harmonic — equally spaced overtones — and so when you take a fourier transform in the frequency domain you'll get a peak corresponding to that equal spacing, which will provide you with the fundamental frequency. (Even if it's missing!)
Using the mel scale totally wrecks that periodicity and throws away the pitch information. (Which is part of why it's used for speech-to-text! In those cases you want to throw away pitch information. Unless you're processing a tonal language, in which case probably don't use the mel scale.)
[+] Why logarithms? In the frequency domain, most harmonic sounds look like the product of a high frequency periodic "signal" (the fundamental and its overtones) with a slowly varying signal (frequencies which are emphasized or de-emphasized, like e.g. formants in the case of speech). Taking the logarithm splits that into the sum of a high frequency signal (overtones) and a low frequency signal (formants), and since the fourier transform is linear, that'll show up as a single peak in the cepstrum corresponding to the gap in overtones (ie the fundamental) and some stuff in the low-frequency bins corresponding to the formants.
The 19hz error being explained by the FFT resolution only makes sense for a classification-based loss/error that used the (FFT_bins / 2) as classes.
Since the proposed network is using regression, even though you have a frequency resolution of 19hz, you should be able to estimate pitch with finer resolution if you are using any popular non-rectangular window because it can be fit to match the shape of the main lobe. You would only expect such a large error at the very low frequencies, where there wasn't much to interpolate on because the next harmonic would overlap.
For an example see figure 4 in PARSHL, (One of the original sinusoidal analysis frameworks where the frequency of each harmonic is estimated by fitting to a parabola) https://ccrma.stanford.edu/~jos/parshl/parshl.pdf
A neural network should be able to do much better than parabolic fitting.
I'll give it a shot (and edit the post.)
mpv and ffmpeg come with CQT visualization:
mpv --lavfi-complex="[aid1]asplit[ao][a]; [a]showcqt[vo]" "$@"
You can even get it from microphone with some piping:
parec --latency-msec=1 | sox -V --buffer 32 -t raw -b 16 -e signed -c 2 -r 44100 - -r 44.1k -b 16 -e signed -c 2 -t wav - | ffplay -fflags nobuffer -f lavfi 'amovie=pipe\\:0,asplit=2[out1][a],[a]showcqt[out0]'
Calculating CQT can be roughly as fast as FFT.
http://academics.wellesley.edu/Physics/brown/pubs/effalgV92P...
And here are some real musical samples you can use instead of the artificial midi notes:
I did though spend a lot of time studying non-DL approaches to pitch-detection, mainly because I wanted better real-time performance for my game Pitchy Ninja (https://pitchy.ninja).
Although I was able to test with real instruments (and my grotesque voice), I didn't find any good live examples of audio with missing fundamentals to test with. It did recognize held-out synthesized data correctly though.
--- Pitch detection (also called fundamental frequency estimation) is not an exact science. What your brain perceives as pitch is a function of lots of different variables, from the physical materials that generate the sounds to your body's physiological structure.
One would presume that you can simply transform a signal to its frequency domain representation, and look at the peak frequencies. This would work for a sine wave, but as soon as you introduce any kind of timbre (e.g., when you sing, or play a note on a guitar), the spectrum is flooded with overtones and harmonic partials. ---
All this said, you don't need deep learning for decent pitch detection -- it's a solved problem, and there are lots of well-known algorithms for it. Deep learning is useful for more advanced music info retrieval such as interval and chord recognition, which was one of my goals with this experiment.
What you will get is a variant spectrogram. Method is related to edge directed interpolation and bilateral transform. (You can do the same with log-cepstrum.)
I find this whole exercise linked an undergrad level toy, even in such basic thing as pitch.
The real problem is polyphony and instrument segmentation which this does not begin to touch.