RNNoise: Learning Noise Suppression
people.xiph.org
people.xiph.org
RNNoise on the other hand seemed to detect silences well, but left artifacts in the speech such that it had a choppy and artificial feel. Lacking the smoothness in the background I found I was more distracted by the distortions in the words.
To me it sounds like the kind of flange-y wafty MP3 glitches you used to get. At 0dB on the babble sample, it's painful to listen to whilst Speex is perfectly fine.
For me, Speex wins on all samples.
At least that's my take on why Speex is subjectively more pleasant, with which I agree.
I was slightly confused by the text that claimed that it was expected for the intelligibility to go down though, so maybe that's all working as intended.
I did a quick hack to RNNoise to smooth out the attenuation and prevent it from cancelling more than 30 dB. I'd be curious if it improves or makes things worse for you (compared to the samples in the demo): https://jmvalin.ca/misc_stuff/rnn_hack1/
With the Speex/raw ones I have all the data so if I listen to it again over and over I can get more out of it eventually.
With the RNNoise one I obviously don't even have enough extra data to even try doing that so all I can do is blame the algorithm.
Perhaps what you really want is an algorithm that lets through a bit more of the 'possible noise' for the human brain to have another go at.
That being said, I still have control over the tradeoffs the algorithm makes by changing the loss function, i.e. how different kinds of mistakes are penalized.
For "car" RNN sounded as good or better than Speex at all noise levels.
That rnn_hack is significantly better for me. 5dB on that sounds strictly better than 10dB on the original to my ear for "babble" and "street". I also noticed that the for the parts that sound the worst to me at 10-15dB in the original RNN, the signal is completely missing in the 0dB RNN version, so perhaps the signal is in the same band as the noise at that part?
Either way it's a tough tradeoff because I suspect that low bitrate encodings will love the nearly empty signal in the bands that are generated by the original, but the seemingly rectangular cutoff/introduction of the noise was much more jarring to me than the reverberation added by Speex (though I didn't like that in Speex, it didn't seem to add to my effort to understand the way that.
I prefer your hacked RNNoise version to the original, but I still prefer the Speex version. Robotic is predictable, and predictable is good. I don't want denoising artifacts to feel like there's some intelligent agent behind them, just as I don't want any software tool to feel intelligent. The smarter the tool the more jarring it is when it misreads my intentions. It might help average performance but it harms worst-case performance, and worst-case performance is subjectively more important because humans pay attention to outliers.
audacity screenshot https://d4344e4d9b25f298d9ea-790118db7dd23376c2de685644429e7...
input https://d4344e4d9b25f298d9ea-790118db7dd23376c2de685644429e7...
RNNoise https://d4344e4d9b25f298d9ea-790118db7dd23376c2de685644429e7...
naive audacity filter https://d4344e4d9b25f298d9ea-790118db7dd23376c2de685644429e7...
To me it sounds like we've got some room for improvement all-around... but color me impressed. I'm also always impressed by Audacity's noise removal when I use it for stupid simple voice-overs. I'd bet this deep learning approach will do nothing but improve, quickly.
$50 will get Audio Cleaning Lab. If I was home I would do a quick auto clean and then hand scrube the file. http://www.magix.com/us/audio-cleaning-lab/detail/
$1199 will get you iZotope’s RX. I actually normally get better results form the $50 one then when I use this tool at a friend's shop. https://www.izotope.com/en/products/repair-and-edit/rx-post-...
Spectral editors are amazing for removing certain sounds and keeping the over all sound intact. This is where we need to move into. Editing at the Spectral level, which will have a much higher CPU overhead.
Here is a great article showing the different hands on techniques for noise removal. https://www.soundonsound.com/techniques/noise-reduction-tool...
By the way, good noise suppression hardware is also comparatively expensive, see for example the Cedar DNS 2 [1]. There could be some business opportunities in that area.
If somebody credible put together a kickstarter for a FOSS equivalent of RX then I'd back the hell out of it. Mostly I think what's needed is a GUI around OSS stuff that already exists, either as plugins or code that can be borrowed.
this demo got me thinking: if I want to remove something very specific from one track instead of learning a generalized filter, can I train this model with a smaller dataset, like a few seconds from that track?
I'd think it would be possible to create something that would do what you're looking for, but it would be much more complex than the above (and -way- beyond what I'm capable of at the moment, maybe in a couple of years I'll be able to do something like it).
I've had more luck with taking the backing and using phasing to remove it from different sections of a song - if you get a track where the backing is simple, sequenced and samples/repeatable synths (so that the sound is identical each time it happens), then it's possible to take that non-vocal section, and align it with the vocal section on another track and reverse its phase to get cancellation; You have to be precise and get lucky in terms of the rest of the track, but it is possible. There is, of course, the old stereo swap and reverse phase trick which removes everything that's not panned centrally; that can get you a lot of mileage.
As mentioned, though, in another comment, getting hold of acappellas/stems can be much better, and having listened to some of some classic tracks, you can learn a lot about production in a short time by doing so.
- the approach is pretty cool!
- as mention in the article, it might be very useful when applied to multiple speakers (conferencing)
- it might be very interesting for speech recognition softwares
Also, as a sound guy, when I have a noisy signal I sometimes remove it a bit too heavily -> I mask the artifacts with some background music. I will definitely try that with the RNNoise suppression !
As strange as it may sound, you should
not be expecting an increase in intelligibility.
I thought one of the reasons hearing aids were so bad was that they pick up noise equally. Wouldn't this method have a direct impact on making hearing aids better?I also have a real hard time differentiating people talking in Google hangouts, say, especially if they're using silverware on porcelain. Wouldn't this type of noise suppression help in this case as well?
Seems like pretty awesome stuff.
I feel this could be taken a step more in such that when the speaker and overlapping loud sound happens at the same time it is able to extract just the speakers voice.
Now obviously this is easier said than done.
That way one could build much better music visualizations programs, and also be a little more creative. I know I have some ideas if I could do it...
I think Funk and related genres are particularly suited for tasks that demand concentration. Funk de-emphasises melody in favor of rhythm. Melody calls for "active" listening. Funk is at the same time predictable and varied. I spend very little time clicking "next track". For me it's very stimulating listening.
So, a problem that could potentially be solved by neuroscience research and programmers (I understand this is interesting in itself) has been solved by good old playlists for me. And experimentation (trying new music).
I disagree with the assessment that if melody calls for active listening that rhythm somehow does not. It really depends on the listener and what they value.
There was a time when I might have agreed with you. That was before I learned to play drums.
Or "Peter" from Holderlin's Traum: https://www.youtube.com/watch?v=iGn6nLaTfqU&feature=youtu.be...
Out of curiosity, could you recommend some bands with very good drums ? One that springs to mind is Guru Guru. When I stop to think about it it seems it's the drums that seal the deal for me, in many tracks. And Embryo, especially on Embryo's Rache.
https://www.youtube.com/watch?v=1yVWmeFM3-Y
https://www.youtube.com/watch?v=GIUlh8qweAc
Something more western - Mastadon - https://www.youtube.com/watch?v=hwgqenxNUfs
Link? :)
Here is an instrumental Funk playlist on Spotify that I have been passively listening to: https://open.spotify.com/user/spotify/playlist/37i9dQZF1DX8f...
I think my adventure started with very generic youtube queries for 'afro' or something like that. Then I stumbled upon Black Merda:
https://www.youtube.com/watch?v=LHSFsWZM1Gk&list=RDLHSFsWZM1... ... and I already liked psychedelic / krautrock.
But it turned out Black Merda is on a very long playlist itself, a playlist I'm disturbingly compatible with !
I'm far from even the half of the playlist, but some of my favorites (other than Black Merda which is awesome): https://www.youtube.com/watch?v=q59ZZtiLgYU&list=RDLHSFsWZM1...
https://www.youtube.com/watch?v=ubyOg-K62co&list=RDLHSFsWZM1...
https://www.youtube.com/watch?v=S6lDGgs7jAc&list=RDLHSFsWZM1...
https://www.youtube.com/watch?v=eaRhetAoqEw
Watermelon Man https://www.youtube.com/watch?v=3FzNpto-jnU&t=1164s
The Variations - Saying It Doing It https://www.youtube.com/watch?v=5OQl4WTkWTc&t=2549s
Only this track, really: https://www.youtube.com/watch?v=-gXrS6eKfjk&t=841s
Plus, there's the joy of hearing silly things like "Back when Dinosaurs ruled the Earth there was a disco not many people knew about." I didn't have so much fun since planet Gong.
I also got some nice recommendations from listening Sisters of Mercy, including Billy Idol. https://www.youtube.com/watch?v=AAZQaYKZMTI And I'm amazed how underrated Max Sedgley is. Must be the last name. https://www.youtube.com/watch?v=ugEgKA24dig Makes me want hone my Inkscape skills.
(What about non-lyrical music? I'm not sure. The research literature on music interfering/non-interfering with cognition is old, large, and highly equivocal in my impression: https://www.gwern.net/Music-distraction )