Coqui, a startup providing open speech tech for everyone
github.com
github.com
For anyone that hasn't heard of Coqui frogs before, they are pretty cool animals. Little guys, but a single one can be surprisingly loud and throw its voice pretty effectively. AFAIK they're only really found in Puerto Rico - apparently they can survive in other warm climates but will not sing? Maybe that's an urban legend though.
Anyway I know the sound is a little contentious (some hotels get cats to cut down on guest complaints about the Coqui) but I'd recommend checking it out: https://musicofnature.com/coqui-magic-nightscapes/
I understand they don't want non-ascii in the GitHub project name that turns into a directory name, but the one-line description contains an emoji so I would have thought they could have allowed themselves a Latin-1 character there, for the benefit of people who know some Spanish but hadn't heard of the frogs.
They tend to quiet down during dry spells, so our rain last week has made things sound much nicer.
Please kill all the Coqui frogs in the world. Or, at least imprison them all on an island that has no human habitation.
Sure, it’s cute in small quantities, but if that’s your soundtrack 24x7 at 70-80 decibels or more, non-stop, you’d probably want to commit suicide just to get out of there. That’s the kind of place that Coqui frogs drive you to.
https://soundcloud.com/user-565970875/pocket-article-wavernn...
Basically I'm wondering if these projects count as libre machine learning projects according to the Debian Deep Learning Team's Machine Learning Policy.
So for the English model they use mostly free/open-source data, but some non-free data:
- train_files Fisher, LibriSpeech, Switchboard, Common Voice English, and approximately 1700 hours of transcribed WAMU (NPR) radio shows explicitly licensed to use as training corpora. (from https://github.com/coqui-ai/STT/releases/tag/v0.9.3)
As for the hardware for training... it's basically NVIDIA (you need Tensorflow / CUDA and all that guff). For inference it works realtime on a CPU.
I'm preparing pretrained models based on Common Voice data for their STT based on Common Voice data here: https://tepozcatl.omnilingo.cc/manifest.html
There is plenty of free/open-source voice data out there, just it's a question of reaching for a sick bag and installing NVIDIAs stuff.
I don't know if they'd count as free/open-source according to Debian (I'm a Debian user myself), but the team has definitely talked about getting into Debian and would be very open to discussions about it.
Using CUDA also means it wouldn't be considered legitimately free enough, although folks are working on getting AMD ROCm into Debian.
Tensorflow isn't yet in Debian, but there may be folks working on it.
Another problem is that Debian doesn't have the hardware for doing training.
I'd encourage you to talk to the folks on the debian-ai mailing list and IRC channel to discuss these and other issues.
https://lists.debian.org/debian-ai/ ircs://irc.oftc.net/debian-ai
If I were into conspiracy theories I'd say that AMDs failure to compete in the GPU/DL space is to do with the relationship between the AMD CEO and the NVIDIA one.
Tensorflow is just awful, as is anything that touches bazel :)
Mind if I ask why?
Fwiw, that sounds like a bug or a misconfiguration; it's absolutely supposed to have better caching behavior than that (and does in the few projects I've used it on, even on a personal laptop). If you're interested in pursuing it further (I'd understand if you aren't; that sounds frustrating), I bet the bazel team would be interested in your report.
[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
Most data is protected by copyright, but I assume you meant proprietary rather than copyrighted. Using proprietary data might not matter under copyright law, but it does matter in terms of the Debian machine learning policy and DFSG, because the non-free data cannot be shipped in Debian main and thus cannot be used to train a model shipped in main.
No doubt about that, but you need validated transcripted voice data(no errors) and this is harder to get.
Aside from Common Voice, there are also a lot of resources at openslr. Also, the amount of data you need is often vastly overestimated, with advances in pretraining and transfer learning and the fact that most languages don't have as terrible an orthography as English.
- Mozilla fired the developers and mothballed the project
- But wants to keep it around as a museum piece
All ongoing development is happening in the fork.
edit: I found some transcription code:
https://github.com/WebThingsIO/voice-addon
DeepSpeech based but close, workable.
If you have any more specific requirements then we can point you in the right direction. Or just join us on Matrix: https://app.element.io/#/room/#coqui-ai_STT:gitter.im :)
And I don't want the company that must not be named to know what I said.
Otherwise this code does streaming on a websocket: https://github.com/coqui-ai/STT-examples/tree/r0.9/web_micro...
I'm looking forward to bothering you on the Matrix ;)