Project Common Voice
voice.mozilla.org
voice.mozilla.org
No matter what, the collected data might be useful for all of that, maybe except of voice recognition actually, because I guess the data will be collected anonymously?
Note that there are some other existing big open speech corpora such as LibriSpeech (http://www.openslr.org/12/) which could already be used right now to build a quite good speech recognition system.
Perhaps they want to do both eventually (?) That could explain the name.
All the text-to-speech software I've used has a generic sounding accent for the country (your choices are typically American, Australian, British, Canadian) but there's a lot more accents out there. The software isn't bad - it sounds realistic - but I wish it sounded more how I would like it to.
There's some software, e.g. Cepstral Dallas - https://www.cepstral.com/en/demos but it sounds too robotic to actually use and that voice isn't available for Linux so I only have it installed on my MacBook.
I guess a lot of developers at e.g. Apple live in CA so Siri is probably influenced by that.
To answer you question: Common Voice is about building a collection of labelled voice data (ie. sentence clips w/ transcripts) that can be used to, for instance, train speech-to-text algorithms. Part of the goals of this project though is to figure out how this data can best help people build voice technology. So it's pretty open ended at this point.
Mozilla does have an open source speech-to-text engine [1] we are developing, and we hope one day to use the Common Voice data to train this engine. DeepSpeech and Common Voice are related, but separate projects, if that makes sense.
As for LibriSpeech, the DeepSpeech team at Mozilla does use this data for training. However, the language is pretty antiquated, and we only get about 1K hours of data, whereas you need about 10K hours to get to a decent accuracy (WER of 10% and below). Common Voice is about adding to public corpora like LibraSpeech, not replacing them.
I would also not use voice technology as the generic term for speech recognition, text-to-speech, and whatever else you want to do with this data. Rather, speech technology is the common term to cover all of this (https://en.wikipedia.org/wiki/Speech_technology).
[0]: An example sentence is "a thin circle of bright metal showed between the top and the bottom of the body of the cylinder", which is from H. G. Wells' War of the Worlds.
This gives Amazon, Apple and Google a nice advantage since they are able to collect huge sample sets of actual voice commands used by people and to some extent also correlate them with the actual action taken by the person.
How could we collect such dataset? It's a bit chicken-egg problem. I don't want to talk to some open source system unless it has fairly good chance of understanding me. Should we try to half manually (through crowd sourcing) come up with potential requests like "Check news from CNN.com", "Order me quattro stagioni" which could be then fed to platform like Common Voice?
Or should we work on higher level. Come up with task descriptions ("You want to order taxi to get to airport for your morning flight at 7am") and then let people record how they would actually request this from computer with voice. This might more accurately capture the language we actually use when speaking. Through some simple automation you could generate variations of the requests and at least partly the same base material could be used for different languages (task given in English, ask person to make the request in Finnish).
I'm wondering if the format will be easily translatable to the kinds of models that software like CMUSphinx and Julius use
If you poke around github and the Kaldi lists a bit more you can see that they are experimenting with and probably planning to use Kaldi.
I wonder what they plan to do for provisioning. It is one thing to collect data and train models, but quite another to make the service available over the web in an unlimited capacity. And we are not yet to the point where you can reasonably expect to run a high quality open-vocabulary STT system in your browser. The search network is typically in the GBs range.
The relevant blurb:
Your Contributions and Release of Rights
By submitting your recordings, you waive all copyrights and
related rights that you may have in them, and you agree to
release the recordings to the public under CC-0. This means
that you agree to waive all rights to the recordings
worldwide under copyright and database law, including moral
and publicity rights and all related and neighboring rights.
[1] https://voice.mozilla.org/termsMost of the computer generated stuff I've seen uses trained actors. Which neatly avoids the problem of trying to reconcile a myriad of accents and dialects, which was immediately apparent from the first two samples I tried.
edit: back up, seems to be about voice recognition, which this could help with no problem.
I think you are correct.
> Read a sentence to help our machine learn how real people speak. Check its work to help it improve. It’s that simple.
> We want the audio quality to reflect the audio quality a speech-to-text engine will see in the wild. Thus, we want variety. This teaches the speech-to-text engine to handle various situations—background talking, car noise, fan noise—without errors.
If you are a non-native speaker, we need your voice!
Anyhow: these should be merged (even though there is no discussion on the other submission)
https://hn.algolia.com/?query=dang%20deliberately%20porous&s...