Mozilla Overhauls Speech-To-Text Contribution Interface
voice.mozilla.org
voice.mozilla.org
It's awesome that the dataset is offered with a CC-0 license: https://voice.mozilla.org/en/data, does anyone know if it includes the answers from the survey? I have a limited bandwidth internet, so I haven't checked it out yet. In particular, I'm wondering if there's user data on whether they picked yes or no, to implement "troll detection" - people who click one option all the time.
Trolling of crowdsourced data isn't unheard of. (NSFW language: https://www.reddit.com/r/pics/comments/cygfx/4chan_is_using_...)
...in that case it was specifically done in retaliation to Google trying to get free mental labour from ReCAPTCHA users in return for being able to post to 4chan. Quite a different situation.
I'm downloading it now, I'll have an answer in a half hour. Does anyone know if there is a torrent for it?
Anyhow, here's a sample from the csv file:
filename,text,up_votes,down_votes,age,gender,accent,duration
cv-valid-test/sample-001224.mp3,but i felt miserable watching him wither away like a shriveled dandelion,1,0,thirties,male,england,
Not sure how some of these are being populated, but yeah; there's several additional folders including invalid mp3, a splintered train set (not sure how it was selected) and a test set folder.Here's the README.txt. Looks cool! Have happy hacky fun! :)
https://gist.github.com/cwgreene/f7f4df4ddcd9da017b9f4694b3f...
Interestingly; many of the 'invalid' mp3's are actually (mostly) correct. Listening to them is interesting to guess as to why they were downvoted.
I see that you have instructions for s3, are the files actually backed in s3? Is it possible to download them with s3 (possibly using requester pays)?
https://discourse.mozilla.org/t/the-mozilla-guarantee-publis...
1.) https://groups.csail.mit.edu/sls/publications/2010/McGraw_LR...
The other thing is that it's very cool to see the "you helped us reach out x% goal" thing but it locks up all the previous / next shortcuts which means I have to switch back to the mouse after 5 entries.
had similar issue/concern. ideally if enough people mark something as correct, the variations and slight differences will get merged together. it did still bother me a bit, as being able to add a bit more extra data would probably be helpful. but... maybe they can add some geo-ip data - respondents from various areas would probably mark more stuff 'correct' from their own region. ???
Being able to mark something 'close', or rate it (1-5, maybe) would help. Just heard an indian accent reading "It's such an unfair world, innit?" The words are... correct, but 'innit' is somewhat idiomatic (especially spelled out that way - seems more UK-oriented text). The pronunciation was "correct" but "awkward".
Also... (too lazy to check right now) - if I create an account, can I see the 'yes/no' ratings of my own submissions?
Not yet, but this is something in the works. You can explore our new experience with the evergreen link: http://bit.ly/cv-desktop-ux
It is something we are still working on.
> The other thing is that it's very cool to see the "you helped us reach out x% goal" thing but it locks up all the previous / next shortcuts which means I have to switch back to the mouse after 5 entries.
That's a bug! Would you mind filing one here: https://github.com/mozilla/voice-web/issues
We have a workaround in place, btw: https://github.com/mozilla/voice-web/issues/1179#issuecommen...
No I shit you not.
Who was it? Miss, present thyself.
The English (in)fluency is more of a feature, though, than a bug. The goal isn't to produce a speech-to-text system that can recognize a perfectly miked BBC announcer. It's to be able to recognize a wide variety of people speaking fairly naturally in imperfect conditions, using whatever accent they use for casual speech.
Wait what? The headline is about text-to-speech aka speech synthesis, not speech recognition (speech-to-text.) Are they trying to do both? It seems to me that you'd train both using different sorts of datasets. If you wanted TTS to be intelligible to the most number of people, training to to speak like a 'perfectly miked BBC announcer' is probably exactly what you'd want to do.
Train it to recognize many regional accents, but train it to speak with the most prevalent and universally understood accent you can find. So either BBC English or Californian/Hollywood English.
Although traditionally TTS engines have shipped with numerous voices, such that you can select either a British or an America accent for the English voice. It may be worthwhile to have other English accents too, maybe one for India (125 million speakers.) But if you trained a TTS engine to have a computer amalgamation of all possible English accents I really doubt the result will be considered high quality by anybody.
For what it's worth, they do ask you to create a profile after your fifth sample, and that profile includes an "accent" section.
It would be kinda interesting to have a TTS system learn from a neighboring STT system so that it gradually adopts your accent, though. I'm not sure if that would be more usable but it would be an interesting experience.
That's exactly what I've been doing recently, and using it with https://github.com/r9y9/deepvoice3_pytorch/blob/master/READM... is providing reasonably good results - it definitely has my intonation (if somewhat crossed with a Dalek!!)
Common Voice is a project to help make voice recognition open to everyone. Now you can donate your voice to help us build an open-source voice database that anyone can use to make innovative apps for devices and the web.
I'll be the first to note that here's another piece of personally identifying information you just "donated"...
Indeed Common Voice is not for everyone. We try to make it clear in our Privacy Policy [1] what pieces of data we collect and why. We do not publish email address or names with the data, and we even strip speaker identification info (so that a speaker's recordings are not grouped but instead everyone's recordings go into one giant bucket). That said, if this still makes you feel uncomfortable, we understand. And if you would like to contribute without donating your voice, you can always validate the recordings of others.
Edit: My bad, is only for the unavailable languages.
But adding new languages to Common Voice is a bit complicated at the moment, and we haven't built a way to do this through the website yet. So for now, we are doing this through a very manual process, and we plan to use email addresses to communicate.
If you want to read more about it, the GitHub repo and in particular the issues cover a lot of the obvious questions like this. They're here: https://github.com/mozilla/voice-web/issues
However, Mech Turk is better for things like validation, since you only need a handful of people doing the majority of work.
In any case, I have some very hacky tools we used for this exploration, if you are interested: https://github.com/mikehenrty/mech-turk/