I noticed that RNNoise doesn't appear to be an open model, you can't re-train it from scratch from the source data, which isn't publicly documented (or doesn't exist?), even if you had enough hardware.
https://www.ntt-at.com/product/multilingual/ https://www.ntt-at.com/product/speech2002/
I don't feel I'm qualified to elaborate on this specifically, again, I'm no ML person. For more info look here: https://github.com/xiph/rnnoise/tree/master/training https://github.com/xiph/rnnoise/blob/master/TRAINING-README
16:57 <ArsenArsen> where and under what license is the training data used for RNNoise?
18:38 <rillian> ArsenArsen: There's a copy of what I believe is the training data on the xiph server, but afaik it's never been published
18:39 <rillian> the original submission page has an EULA waiving copyright and liability claims, and agreeing that it _may_ be released CC0.
18:40 <rillian> it looks like that didn't actually happen.
18:41 <rillian> there may have been concerns about auditing it for privacy issues, but there's a lot of audio to listen to, 6.5G compressed
18:41 <rillian> jmspeex, TD-Linux: what's the status of publishing the rnnoise training data?
18:43 <jmspeex> Are you talking about the data that was used to train the default RNNoise model or the noise that got collected with the demo?
18:43 <rillian> jmspeex: I think debian just cares about the training data for the default model.
18:44 <jmspeex> There was never plan to release that -- it includes data from databases we cannot release
18:44 <jmspeex> but I don't see what the issue is. Distributing the model is not the same as distributing the data
18:45 <rillian> ah, I see. I didn't realize you'd used proprietary sources as well.