Enhanced noise suppression in Jitsi Meet
jitsi.org
jitsi.org
Overview of RNNoise from the horse's mouth is here: https://jmvalin.ca/demo/rnnoise/
Used as a Wasm module! In some ways the web is becoming more opaque. Is this the future then, a hodgepodge of binaries doing things behind the scenes? Though in this case it happens to be OSS, and it may well be a moot point -- backend is already a blackbox to the enduser, now parts of frontend are blackboxes. The practical implication is probably just that some measure of customizability is gone.
Compiled code has existed for half a century and we know how to work with it.
Suggesting that the web is doomed because people of the future prefer rust instead of javascript is beyond any rationale.
The early web was a great equalizer. Anybody could study a little html, download an ftp manager, jump through a few procedural hoops and have a web page. After some studying and trial and error they could even build an interactive site.[1]
It's easy to miss all the potential of wasm when that's what you remember of the web. To me the amazing thing is that browsers will still work with the methods described above[2] but we're on the cusp of being able to do almost everything a full application environment can do.
That said, even though there will be plenty of OSS wasm tech, it'll still be more opaque to those of us who don't do compiled languages. It'll be a lot tougher to just fork the code and do something more creative with it.
[1] PHP used to stand for "Personal Home Page" and, as one of its founders put it, was created so that "any idtiot" could make an interactive site.
We already lost any semblence of building from scratch in the mid-2000s with the emergence of gargantuan HTML templates and Wordpress/Drupal/PHPbb deployments with plugins and themes.
This is a direct result of people being held to higher standards and thus spending a lot more effort overriding the compositional and behaviour defaults of the user agent.
The modern-day iteration just optimizes for scaling up to tens of thousands of concurrent end-users on anemic hardware.
We have to accept the fact that personal webpages gave way to social network profile pages. This didn't happen overnight and there is zero demand for a hand-crafted presence on the web anymore.
I dont think "no code" is an aid. If anything it's pushing in the opposite direction: rather than a transparent approachable web medium, it suggests we need hyperadvanced tools that we really wont understand or have control over to synthesize web code. It's a simpler user experience, but a push away from notepad.exe webdev.
I wouldnt rush to make any conclusions about who or what has won, as a settled fact & case for all time. We havent had good ways to run online systems ourselves, versus hosted for us, and there's still lightyears to go but we're doing good things & finally maturing well. We're only a couple years into ActivityPub as an interchange format & growing many of the caoabilities & tools & systems, around all mimds of use cases, that will make throwong together a fair, interactabke competitive offering possoble. Social media has had huge huge investmemt poured into it, but we are in decent preteen years of growing up & owning the libre equivalents. We can assess demamd only after there is a visualizable state people can imagine; just having an isolated blog is not the equivalent to the well connected social media site, but these capabilities slowly arise. Follow the alpha geeks; this currently long phase will not be forever.
When I was a kid, every piece of software I used was pre-compiled, and therefore opaque. This made it difficult for me to figure out how people made certain things, and after a while I lost interest in programming.
When I got back into it later, one thing that made a huge difference was being able to see how various cool JS sites were built. The ability to “View Source” like that was revolutionary, and also allowed me to build some early fun projects, like a Cookie Clicker “AI” that could play the game automatically by calling the functions I could see in the game’s source.
I’m far from the only person with experiences like these. Yes, there was programming before View Source and there will be programming after. And for those of us with the right tools or reverse engineering skills, View Source isn’t particularly relevant. What we’re losing is a pipeline that helped people become/stay interested in programming, which makes it likely that future programmers who would’ve followed a path like mine will do something else instead.
If that’s your definition of transparency, then perhaps learning to read assembly would give you the same comfort. In fact, there’s a lot more binaries distributed with symbols intact than unminified JS.
Or, to put it another way, if you could right click -> view disassembly of any binary on your computer, how would that be any different than today’s web?
(You could technically turn the wasm to JS and unminify that too, which I doubt is much harder/easier to decipher as the same thing written in JS and minified/unminified.)
How else would we have implemented this? WASM has facilitated introducing these technologies into web applications, it literally wasn’t possible before.
Thanks to emscripten it wasn’t even that hard to get rnnoise working on WASM: https://github.com/jitsi/rnnoise-wasm
I concede WASM does open the possibility of adding opaque stuff to web apps but IMHO the benefits outweigh the drawbacks at this point.
When I install a program on my debian machine with apt-get, I also get binaries. But this doesn't mean that it is opaque right?
I didnt understand the first three words, for Alice it was the next two, and for Bob it was the last four. How many people are going to ask to repeat?
Evolution taught us to understand over the sound of waves, crickets, rain, thunder, and more. It didn’t teach us to comprehend with half the signals masked.
> So what should you listen for anyway? As strange as it may sound, you should not be expecting an increase in intelligibility. Humans are so good at understanding speech in noise that an enhancement algorithm — especially one that isn't allowed to look ahead of the speech it's denoising — can only destroy information. So why are we doing this in the first place? For quality. The enhanced speech is much less annoying to listen to and likely causes less listener fatigue.
> Actually, there are still a few cases where it can actually help intelligibility. The first is videoconferencing, when multiple speakers are being mixed together. For that application, noise suppression prevents the noise from all the inactive speakers from being mixed in with the active speaker, improving both quality and intelligibility. A second case is when the speech goes through a low bitrate codec. Those tend to degrade noisy speech more than clean speech, so removing the noise allows the codec to do a better job.
I do think that for direct listening, the jitsi.org speech samples would be slightly more intelligible if the noise removal was tuned to pass through frequencies with mixed noise and signal. I don't know if that would be worse in a video conference. Does the speaker or listener get to choose between conservative and aggressive noise removal?
Nice to see it getting integrated into video meeting solutions, so more people can take advantage of this awesome library.
Have you tried https://github.com/noisetorch/NoiseTorch/?
https://github.com/werman/noise-suppression-for-voice#pipewi...
I'd love to see some more state of the art solution that works with WASM. Maybe even something that one could train on their own voice and filter everything else would be awesome. Because all the noise cancellation tech does not help if you sit in an environment with other people talking next to you and the AI doesn't filter it because it's voices. Sometimes coworkers use Krisp but even that proprietary paid solution is so-so.
Also audio worklets weren’t a thing when we first introduced it.
I’m not aware of any other open source (and better) models, but if any come up, we’ll certainly check them out!
On that topic — they record sessions in an interesting way, basically an instance of chrome is started and captured... I think with OBS. That always made me raise an eye but I also can’t think of up a better way.
edit: It's actually jibri which has to do with recording. Gosh I wish the names were a liiiittle more intuitive. :)
At the time I also spent a few days looking for something better but didn't really find anything. Unfortunately RRNoise is the best we have :( The only other noise cancellation software that actually impressed me was the one from Nvidia but that's not something that one could integrate via WASM and of course wouldn't work on most devices anyways.
Oh what a day it will be where we have energy efficient hardware encoders for AV1 in every device plus some really good noise cancellation. Oh and then we just need internet connections without packetloss :P
Jitsi Meet has been a great alternative to other meeting apps in these crazy times.
I have run jitsi on cheap VMs and it worked decently. But you need quite some cores to serve all the traffic. Ultimately I ended up having as many 2-4core VMs as I had concurrent calls.
1: https://support.apple.com/guide/iphone/change-the-audio-sett...
Does it create friction for folks who haven't used it before? Any suggested instructions to send with a meeting invite?
We do unfortunately see semi-regular lock-up/freezes where one end of the stream stops for ~30 seconds. Maybe this is worse in safari vs chrome/Firefox - we have not yet experimented much with different browsers. Or maybe there's a difference between x86_64 and arm/m1/m2.
- https://bugs.webkit.org/show_bug.cgi?id=244203 - https://bugs.webkit.org/show_bug.cgi?id=241223 - https://bugs.webkit.org/show_bug.cgi?id=221334
Personally, I'd stick with the big names, long remote meetings are strenuous enough even with all the quality of life features those offer.
Most devices have noise suppression built in. Do you use RNNoise in tandem with the hardware noise suppression, or do you disable it?
Also, are there any plans to implement this with react-native on mobile? And if so, how would you implement this, since audio worklets and WebAudio aren't available.
A desktop client also exists for Windows, macOS, Linux: https://github.com/jitsi/jitsi-meet-electron - kind of not really advertised, provides remote desktop control contrary to the strictly web browser version.