Would it be possible to upload a few before / after samples with varying degrees of background noise? Even if it's all the same person that would be a huge help to gauge the quality.
I just wont get to it today unfortunately.
Just a suggestion if you do it, please include realistic room noises in some of the samples.
I looked at the RNNoise examples and it was pretty bad. I mean, the audio quality of the speaker got completely mangled but the background noise was also comically high. It sounded like the person just sat down in the middle of the street in NYC or was inside of a busy train terminal.
This works really well for situations like Discord or Teamspeak where you're usually not constantly talking, but doing things that can still set off "normal" voice activation. RNNoise's model often knows it's not voice, but cannot denoise it completely.