HNHacker News
TopNewBestAskShowJobs

punchingwater

135 karma · joined December 30, 2015

Hi, I'm Michael. You can find more of me here:

https://github.com/mikehenrty

and here:

https://twitter.com/mikehenrty

submissionscomments
punchingwater··on Mozilla releases the largest to-date public domain transcribed voice dataset
Audiobooks are definitely possible for ASR training. Indeed the largest open ASR training dataset before Common Voice was LibriSpeech (http://www.openslr.org/12/). Also note, the first release of Mozilla's DeepSpeech models were trained and tested with LibriSpeech: https://hacks.mozilla.org/2017/11/a-journey-to-10-word-error...

But as others have mentioned, there are several problems with audiobooks as an ASR training dataset. First, the language used in literature is often very different from how people actually speak, especially if that language comes from very old texts (which many public domain books are indeed quite old).

Then there is the sound profile, which includes background noise, quality of microphone, speakers distance to device, etc. For recorded audio books, the speaker is often using a somewhat sophisticated setup to make the audio quality as clean as possible. This type of setup is obviously unusual when people want to speak to their devices.

Third, the tone and cadence of read speech is different than that of spontaneous speech (the Common Voice dataset also has this problem, but they are coming up with ideas on how to prompt for spontaneous speech too).

But the goal of Common Voice was never to replace LibreSpeech or other open datasets (like TED talks) as training sets, but rather to compliment them. You mention transfer learning. That is indeed possible. But it's also possible to simply put several datasets together and train on all of them from scratch. That is what Mozilla's DeepSpeech team has been doing since the beginning (you can read the above hacks blog post from Reuben Morais for more context there).

punchingwater··on Danish physicists claim to cast doubt on detection of gravitational waves
the article states:

> “But Reitze counters that the complete data from that first run is already available online. According to Shoemaker, this includes the relevant time series data and the programs used, but "it's not a trivial matter to use them." Caltech even held a training workshop on how to deal with gravitational-wave data. That's a pretty far cry from asking the physics community to take its analysis on faith...”

Are you refuting this statement? To me, it looks like the data/analysis are open, but not yet independently verified due to difficult nature of problem space.

punchingwater··on Speech-to-Text Benchmark: Framework for benchmarking speech-to-text engines
Just to add my two cents (I work for Mozilla on Common Voice): without help from linguists, Common Voice would have made some very different and very bad decisions about all sorts of things like: accents, dialect segmentation, corpus curation, licensing, and many other things. Linguistics were absolutely instrumental. We tried to thank some of them at the end of our blog post: https://medium.com/mozilla-open-innovation/more-common-voice...
punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
We have no plans to allow users to download the "raw" data from s3 (ie. before we perform the train/dev/test split). But we want to eventually build some tools to automate this. See here for some background:

https://discourse.mozilla.org/t/the-mozilla-guarantee-publis...

punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
> The other thing is that it's very cool to see the "you helped us reach out x% goal" thing but it locks up all the previous / next shortcuts which means I have to switch back to the mouse after 5 entries.

We have a workaround in place, btw: https://github.com/mozilla/voice-web/issues/1179#issuecommen...

punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
Just to note, we will never require your email address to contribute. There will always be an anonymous contribution workflow.

But adding new languages to Common Voice is a bit complicated at the moment, and we haven't built a way to do this through the website yet. So for now, we are doing this through a very manual process, and we plan to use email addresses to communicate.

punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
Thank you for bringing this up.

Indeed Common Voice is not for everyone. We try to make it clear in our Privacy Policy [1] what pieces of data we collect and why. We do not publish email address or names with the data, and we even strip speaker identification info (so that a speaker's recordings are not grouped but instead everyone's recordings go into one giant bucket). That said, if this still makes you feel uncomfortable, we understand. And if you would like to contribute without donating your voice, you can always validate the recordings of others.

1.) https://voice.mozilla.org/en/privacy

punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
In the early days of this project, before we shipped the website (ie. ~March of 2017), we did some explorations around Mechanical Turk. The problem with the Mech Turk approach is that for recording your voices you need a lot of different people speaking (ie. 10s of thousands). But for languages other than English, Mech Turk simply doesn't have these kind of numbers. And indeed English is not that interesting to us, since there exists public data already in English (see LibriSpeech). There are of course other micro-task platforms popular in other countries (for instance, there's a myriad in Indonesia), but we didn't have the time to manage jobs on all these different platforms.

However, Mech Turk is better for things like validation, since you only need a handful of people doing the majority of work.

In any case, I have some very hacky tools we used for this exploration, if you are interested: https://github.com/mikehenrty/mech-turk/

punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
There has been some discussion around this, but no real movement yet: https://github.com/mozilla/voice-web/issues/336
punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
Would you mind filing an issue? https://github.com/mozilla/voice-web/issues
punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
> Also... (too lazy to check right now) - if I create an account, can I see the 'yes/no' ratings of my own submissions?

Not yet, but this is something in the works. You can explore our new experience with the evergreen link: http://bit.ly/cv-desktop-ux

punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
We do have a issue filed to allow users to tag recordings with certain metadata, like noisy or male/female voice. https://github.com/mozilla/voice-web/issues/814

It is something we are still working on.

> The other thing is that it's very cool to see the "you helped us reach out x% goal" thing but it locks up all the previous / next shortcuts which means I have to switch back to the mouse after 5 entries.

That's a bug! Would you mind filing one here: https://github.com/mozilla/voice-web/issues

punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
We also keep the README in the repo: https://github.com/mozilla/voice-web/blob/master/docs/corpus...
punchingwater··on Mozilla Overhauls Speech-To-Text Contribution Interface
We used some of the research around Mechanical Turk to find best practices for limiting trolling (e.g. [1]). Our approach thus far has been the two-thirds rule: if two out of three people say the clip is good/bad, we trust that. Also note, that we have seen remarkable low trolling numbers, and that most of the invalid are pronunciation mistakes or saying the wrong word.

1.) https://groups.csail.mit.edu/sls/publications/2010/McGraw_LR...

punchingwater··on Initial Release of Mozilla’s Open Source Speech Recognition Model and Voice Data
Don't forget Kaldi!

https://github.com/kaldi-asr/kaldi

punchingwater··on Initial Release of Mozilla’s Open Source Speech Recognition Model and Voice Data
Thank you so much!

I also want to emphasize the importance of listening (validating) as well as recording. Validation is an big part of the puzzle for building machine learning viable data.

punchingwater··on Initial Release of Mozilla’s Open Source Speech Recognition Model and Voice Data
Yup, this is an excellent point. We have, and will continue to explore ways to allow Common Voice users to speak more organically (for instance by answering a question, or responding free-form to some other sort of prompt). The problem with this approach is that it requires an extra step, transcription, which at the scale we are trying to achieve is pretty costly in either money or time (ie. tedium for our users). Eventually we hope that speech engines can take care of the transcription part, but for now we need people.

That said, we will definitely be exploring ways to build in organic speech and perhaps transcriptions to the Common Voice app. This will solve another problem for us too, which is getting public domain material for people to read. Doing this obviously requires a much more complex user experience, and we have more work to figure out how to make something that people will want to use and contribute to. Stay tuned for that :)

On the flip side, we hope that these datasets, models, and the tools (ie. DeepSpeech) can get more people (researchers, start-ups, hobbyist) over the hump of building an MVP of something useful in voice. Once you have people using your products, collecting useful in-context voice data becomes much easier.

On that note, another approach we are working on is partnering with universities and socially-aware startups like MyCroft, SNIPS, and Mythic. Imagine if voice products in market allowed their users to opt-in to contributing their utterances to an open resource similar to Common Voice. Of course, sharing your voice publicly is not for everyone, or every product scenario. But it does work for some. And if we pool our resources, our hope is to indeed commoditize speech-to-text so that we can focus on more interesting challenges like building voice experiences people want to use. (For instance, could voice somehow be a "progressive enhancement" to the web?).

punchingwater··on Project Common Voice
Noted. Again thanks for the feedback :)
punchingwater··on Project Common Voice
This is a bug with our website [1]. We actually are trying to collect non-native speakers (as well as native). We are looking into clarifying this on the site.

1.) https://github.com/mozilla/voice-web/issues/242

punchingwater··on Project Common Voice
Good question. Sounds like we should add an "Other" to that drop down, and make it clear that we are looking for all accents?
punchingwater··on Project Common Voice
Exactly! Part of the goals of Common Voice is to make voice recognition work better for non-north american men (which is where the vast majority of the training data comes from).

If you are a non-native speaker, we need your voice!

punchingwater··on Project Common Voice
Sorry about the 503s! We were adding servers to our cluster to handle the hacker news load, and a few 503s are hard to avoid. If this is consistently happening for you, please file a bug and we'll look at it.

https://github.com/mozilla/voice-web/issues

punchingwater··on Project Common Voice
Great feedback, we can look into clarifying on our homepage that our entire goal is to create a dataset in the public domain. We want people to donate not just to Mozilla, but to the world :)
punchingwater··on Project Common Voice
Common Voice is only about collecting a large public database of voices. We do have a separate project around speech-to-text [1]. We haven't done much work around speaker recognition (AFAIK) or voice synthesis, but they are both very interesting both from a technical and privacy related standpoint. That said, both are out of the scope of Common Voice (which is only about the data).

1.) https://github.com/mozilla/DeepSpeech

punchingwater··on Project Common Voice
thanks for the vote of confidence! yes we will absolutely open this data up, and it's just a matter of collecting enough data to be useful, and then building the UI. we have a goal of achieving this by the end of 2017, so stay tuned!
punchingwater··on Project Common Voice
I can tell from your comment (and it's responses) that the language on our homepage is a bit confusing, so thank you for the feedback.

To answer you question: Common Voice is about building a collection of labelled voice data (ie. sentence clips w/ transcripts) that can be used to, for instance, train speech-to-text algorithms. Part of the goals of this project though is to figure out how this data can best help people build voice technology. So it's pretty open ended at this point.

Mozilla does have an open source speech-to-text engine [1] we are developing, and we hope one day to use the Common Voice data to train this engine. DeepSpeech and Common Voice are related, but separate projects, if that makes sense.

As for LibriSpeech, the DeepSpeech team at Mozilla does use this data for training. However, the language is pretty antiquated, and we only get about 1K hours of data, whereas you need about 10K hours to get to a decent accuracy (WER of 10% and below). Common Voice is about adding to public corpora like LibraSpeech, not replacing them.

1.) https://github.com/mozilla/DeepSpeech

punchingwater··on In Memoriam: Ian Murdock
> (quote from imurdock's twitter) Maybe my suicide at this, you now, a successful business man, not a NIGGER, will finally bring some attention to this very serious issue.

This sounds a lot more like 4chan trolls than it does one of the leading lights of OSS. I know people are trying not to speculate, but I wouldn't be surprised if his account was hacked. In any case, I think you can honor they memory of Ian and still try to find out what happened to him.