RNNoise on the other hand seemed to detect silences well, but left artifacts in the speech such that it had a choppy and artificial feel. Lacking the smoothness in the background I found I was more distracted by the distortions in the words.
RNNoise on the other hand seemed to detect silences well, but left artifacts in the speech such that it had a choppy and artificial feel. Lacking the smoothness in the background I found I was more distracted by the distortions in the words.
I did a quick hack to RNNoise to smooth out the attenuation and prevent it from cancelling more than 30 dB. I'd be curious if it improves or makes things worse for you (compared to the samples in the demo): https://jmvalin.ca/misc_stuff/rnn_hack1/
With the Speex/raw ones I have all the data so if I listen to it again over and over I can get more out of it eventually.
With the RNNoise one I obviously don't even have enough extra data to even try doing that so all I can do is blame the algorithm.
Perhaps what you really want is an algorithm that lets through a bit more of the 'possible noise' for the human brain to have another go at.
That being said, I still have control over the tradeoffs the algorithm makes by changing the loss function, i.e. how different kinds of mistakes are penalized.
For "car" RNN sounded as good or better than Speex at all noise levels.
That rnn_hack is significantly better for me. 5dB on that sounds strictly better than 10dB on the original to my ear for "babble" and "street". I also noticed that the for the parts that sound the worst to me at 10-15dB in the original RNN, the signal is completely missing in the 0dB RNN version, so perhaps the signal is in the same band as the noise at that part?
Either way it's a tough tradeoff because I suspect that low bitrate encodings will love the nearly empty signal in the bands that are generated by the original, but the seemingly rectangular cutoff/introduction of the noise was much more jarring to me than the reverberation added by Speex (though I didn't like that in Speex, it didn't seem to add to my effort to understand the way that.
I prefer your hacked RNNoise version to the original, but I still prefer the Speex version. Robotic is predictable, and predictable is good. I don't want denoising artifacts to feel like there's some intelligent agent behind them, just as I don't want any software tool to feel intelligent. The smarter the tool the more jarring it is when it misreads my intentions. It might help average performance but it harms worst-case performance, and worst-case performance is subjectively more important because humans pay attention to outliers.
To me it sounds like the kind of flange-y wafty MP3 glitches you used to get. At 0dB on the babble sample, it's painful to listen to whilst Speex is perfectly fine.
For me, Speex wins on all samples.
At least that's my take on why Speex is subjectively more pleasant, with which I agree.
I was slightly confused by the text that claimed that it was expected for the intelligibility to go down though, so maybe that's all working as intended.