Mozilla Common Voice Adds 16 New Languages and 4,600 New Hours of Speech
foundation.mozilla.org
foundation.mozilla.org
I know the Icelandic government spends some money for this and it shows. This tiny language has way more support then other way more spoken languages. If the Norwegian government wanted I bet the Sámi languages could have just as good of a support as Icelandic. Or if the Greenlandic government had more funds available I bet we would see Kalaallisut in more places online.
Serbian government would certainly support Serbian language voice recognition and synthesis, but probably not with as much money as Iceland would.
The idea that there is a difference between these two things is one of the more pernicious ones of the last hundred years.
Money is power. The exercise of power is politics. They can't be separated.
It certainly sounds like this is a political situation to me, almost to a tautology. The fact that these decisions was made on the basis of financial gain doesn't make them any less political.
The reason why there isn't a huge consumer market for indigenous languages is because they're overwhelmingly systematically unsupported by their respective governments in favor of the non-indigenous colonial languages.
To be clear, that's not Mozilla's fault, and not something they or other random organizations can fix, but as human beings we should all be happy and give credit to those organizations that do their small part.
Yet I have never seen an application that had a Sorbian translation, for one simple reason: Every Sorb also speaks German, so there is no financial incentive to invest in a Sorbian translation. The only things with Sorbian translations are those produced in the local area, e.g. the websites of local governments or local businesses.
If the UK is a country of countries then Greenland is most certainly an autonomous country. The Wikipedia article for Greenland has the word “country” mentioned 14 times, so I’m certainly not the only using the term this way.
Wow, I had no idea
The problem is not the Spanish language. The problem is a colonial peasant economy/society that turned into a post-colonial peasant society, with land-owning quasi-nobility ruling over disempowered (in this case, indigenous) laborers and freely exercising their power to steal, rape, kill, etc., without penalty; it is a situation more or less comparable to peasant societies around the world and throughout history, which are always very exploitative and often racist.
Working class Spanish speakers living in towns were in many ways also economically exploited, but considered “better than indigenous people” to be a core part of their identities, and also felt free to beat them, steal from them, etc. where they found the opportunity. It’s a situation broadly comparable to race relations in the US south, where poor whites considered “better than blacks” to be a defining part of their identity.
Perhaps counterintuitively, the history of exploitation of indigenous communities, and the way indigenous people were shut out of many social and economic activities, led to the preservation of native languages.
However I at the same time I’m also deeply disappointed by the lack of support for Iceland’s closest neighbour’s language—Greenlandic—which is an indigenous language, the sole official language of an autonomous country.
Generally, it takes only a few dedicated people to get software localised if good enough infrastructure is provided by the community!
I hope that's what we see with Mozilla Common Voice too!
Luckily, speech recognition research is making some good progress on dealing with low-resource languages so hopefully we'll see some acceptable models made from the little available open data that's out there.
I'm not sure "autonomous country" is an accurate description of what Greenland is. It is - for all intents and purposes - a devolved region of Denmark. It is still way too reliant on economic aid to be able to be independent and, honestly, probably couldn't exist as a developed nation without a patron (Denmark) or without selling its land/resources to a great power (USA, China). And the population is only 1/6 the size of Iceland's and is very dispersed on a massive arctic island, with most people living in tiny isolated villages by the coast.
With that in mind, you wouldn't expect great language support unless the Danish state steps in and spends some serious dough on it. I actually work on Danish language technology at the University of Copenhagen and let me tell you something... the Danish state hardly spends any money on Danish language resources either. We envy the kind of funding that researchers in countries like Iceland and Norway have access too.
I’m actually a little disappointed that there is not more collaboration between the language departments in Iceland and Greenland. Iceland does spend some money on foreign languages and there is much interest in general for foreign languages in Iceland. The former president Vigdís Finnbogadóttir is a huge language buff and advocates for foreign languages a lot. So much so that the house of foreign languages at the University is named after her (https://vigdis.hi.is/).
It is generally believed in Iceland that setting up Icelandic cultural institutions in Reykjavík played a big part in our independence. Institutions such as the University, libraries and the National Theater. There is also big interest for Greenlandic independence in Iceland. Therefor it would make sense for a rich country like Iceland to spend some money in progressing the status of Kalaallisut, both in Iceland (by shared cultural events), Greenland (by help funding cultural institutions) and internationally (by help funding online language efforts).
I don’t think it is wrong to call Greenland a country. As mentioned elsewhere, the word country is not strictly defined. Sometimes it means strictly independent nations, but most of the time it doesn’t. E.g. here is CIA calling Greenland a country (https://www.cia.gov/the-world-factbook/countries/greenland/).
https://venturebeat.com/2021/04/12/mozilla-winds-down-deepsp...
https://blog.mozilla.org/en/mozilla/mozilla-partners-with-nv...
$1.5mil for shutting down open source initiative, almost half of CEO salary right there.
It's significantly closer to "nonfree" on the free-nonfree spectrum than it should be, and is another example of the difference between the guiding philosophies behind "free software" and "open source"
I think real-time factors smaller than 1 are faster than real-time (not slower) and use less than 100% of a resource's computational power to keep up.
> I think real-time factors smaller than 1 are faster than real-time (not slower) and use less than 100% of a resource's computational power to keep up.
Sure, but who has the necessary GPUs installed? And on CPUs it will apparently take longer to generate speech than the duration of that speech. Unusable for many UIs and it will also drain the batteries of any portable device.
So it is the contrary
Especially for the videos with Close Caption....
As simple as extracting the Audio and CC text?
(See https://support.google.com/youtube/answer/2797468 and the part about status.license here: https://developers.google.com/youtube/v3/docs/videos)
Anyone can use the Common Voice data within the terms of the license and NVIDIA contributing towards the continued gathering of data (that will continue to be made publicly available) won't change that.
It's a huge shame that Mozilla didn't continue the DeepSpeech project but Coqui is taking on the mantle there and there are plenty of others working on open source solutions too, all whilst the existence of CV will make a big difference to research, in the academic, commercial and open source spheres.
If that was true that would be a profoundly bad purchase for NVidia since the data is already freely licensed and available for anyone to use at no cost.
This is like saying that Epic "bought" Blender when they gave it a development grant, or that Google contributing patches to upstream Linux means they own it now. Mozilla didn't give NVidia any kind of special license, when NVidia contributes data to Common Voice they're doing so under Common Voice's license, not their own.
We want to encourage more companies to treat software and training data as a public commons that is collectively maintained, this is a good thing.
https://techreport.com/news/14707/ubisoft-comments-on-assass...
https://techreport.com/review/21404/crysis-2-tessellation-to...
https://arstechnica.com/gaming/2015/05/amd-says-nvidias-game...
Here it appears they purchased this https://venturebeat.com/2021/04/12/mozilla-winds-down-deepsp...
And the assumption the shutting down Deep Speech was specifically for NVidia's benefit seems like a fairly large leap to me, given that Deep Speech is already mature, still being developed under Coqui.ai, and surrounded by a wide diversity of other deep learning projects that also aren't controlled by NVidia.
Decreasing barriers of entry for those models and providing raw data is probably the right thing for Mozilla to be focusing on right now. Any team can build a language model, only companies like Mozilla can coordinate mass data collection for those models.
https://commonvoice.mozilla.org/
There are many languages available to pick from.
Most of them were definitely speaking English, but in an Indian intonation that I was barely able to understand coming from an English as a First Language country.
Some of them were reading words syllable by syllable, which is definitely English, but I would hate to have to listen to an ebook or webpage read aloud to me in that manner.
By clicking yes am I training the system to speak English with an Indian intonation?
Should I click no, not English?
Should/does english even have a "proper" intonation?
If your ML model can't handle multiple accents, it is worthless.
Fortunately, splitting models into separate accent-specialized variants and helping them out with language model training will often help in case the model doesn't cope well enough with the cognitive dissonance.
Which american accent?
> Varying Pronunciations
> Be cautious before rejecting a clip on the ground that the reader has mispronounced a word, has put the stress in the wrong place, or has apparently ignored a question mark. There are a wide variety of pronunciations in use around the world, some of which you may not have heard in your local community. Please provide a margin of appreciation for those who may speak differently from you.
> On the other hand, if you think that the reader has probably never come across the word before, and is simply making an incorrect guess at the pronunciation, please reject. If you are unsure, use the skip button.
So don't worry about weird intonation as long as they correctly pronounce the sentences, that way even more people can enjoy the fruit of this labor.
Text to speech should work correctly. But speech recognition should tolerate even clear mistakes. Of course not for the price of misunderstanding correct pronunciation.
Some unusual suspects among the top languages, there!
Esperanto was designed to be easy to learn. It isn't an elite pursuit in the way you suggest, because its community isn't gatekept. I personally have met people of all social classes who have been interested in it.
It was also never meant to be a first language, it is an auxiliary language. It is possible for an English speaker to have a conversation with a Mandarin speaker with no intermediary if both know the (comparatively easy to learn) Esperanto. Its original purpose wasn't trivial either: it was created to stop groups without a common language in the same city (Warsaw, I think?) fighting, created on the basis that they'd stop doing so if only they could speak a common language.
Think of it as JVM bytecode for people.
And that's how it's played out. Nearly every developed nation teaches English as a second language or is a native population of English speakers. The universal language is English. The JVM bytecode for people is English.
What are you telling me? That I need to drop English?
Certainly you will find people learning other languages for trade depending on the region, but even in East Asia, as you say, English is taught in China, Japanese, Korea. In Singapore English is the language everyone learns (and is taught in). In Vietnam the primary foreign language taught is English. In the Philippines one of its official languages is English. Argentina teaches English in elementary school. In Brazil students from grade 6 have to learn a language, which is usually English. In Venezuela English is taught from age 5.
So what exactly do I have to tell them?
I wonder what gave you such an impression of Esperanto. My personal experience of Esperanto is quite different.
I started to casually self-learn Esperanto about one year ago as my second foreign language apart from English. After about half a year, I was confident enough to join online Esperanto communities and it gave me a surprisingly much more diverse experience than any community I had encountered on the Internet.
For example, in an online chat group, active users mainly come from US, South America, and Russia. As an person from East Asia, there is little chance for me to get in touch with the latter two groups otherwise. And there are often new users from South America who speak only Spanish and Esperanto.
I myself do not identify as a upper-middle class person, and I don't know enough to assess other Esperanto speakers' class status.
The impression of Esperanto speakers being upper-middle class may come from the fact people learn Esperanto as a hobby. But people not in the upper-middle class can have other hobbies, why is Esperanto different? It doesn't come with the many benefits that people may expect from learning a "practical" language, but it takes significantly less effort. I'd say it's about as hard as learning a new instrument. So it is not that exclusive to only upper-middle class people.
After one year of casual learning, I am now able to contribute to the Common Voice project in Esperanto (175 recordings and 123 validations) and I actually use it as a source of learning material.
Of course, it takes time to fluently "read out" the words, and in practice, it's much easier if you just know the word and pull the pronunciation from your memory.
For the Common Voice project, there are usually two or three words in a batch of five sentences that I don't know. And there are unfamiliar places and names, since most of the text come from Wikipedia. In such case, I'll take my time to use the spelling to infer the correct pronunciation and practice it several times, until I can put it into the sentence. Then I'll record. And I know it must be correct.
If I am not sure about the meaning of the new word (you can usually guess from etymology or word formation), I look it up in the dictionary and learn a new word.
The point is: the only reason Bengali Korean and Malayalam are stuck "in progress" is that no one is working on them. No language but English is actively supported by Mozilla, it all comes from the communities. And the success of Esperanto shows that every language can make it. I hope that people take our work as a motivation. Every language can become big if a few motivated people work on it for a year or two. Even the smallest language can make it. You just need a lot of public domain sentences, a few thousand donors and some technical knowledge then your language will grow as well :)
When I can use Google or Facebook in any of these languages for 10+ years, it's silly of this project to claim some high moral ground when you can't support some of the most widely spoken languages in the world and stick to languages that hipsters in San Francisco think is cool.
When I did not see my own language in the list a year ago, and I had no clue how to get it there, I reached out to my university contacts that I know used to translate Firefox years ago.
With their help we quickly translated the whole common voice site (it was a prerequisite to start contributing a language) and provided first sets of text to start contributing.
In about a week we started contributing voice for a new language. The Common Voice project is awesome and very well made.
One problem is that data for speech recognition needs to be extremely accurate (i.e. the speech matches the transcript perfectly) and the human review process is infallible and there are quite a number of bad clips that made it past the review process (to be fair, Mozilla provides no official guidance to reviewers or recorders).
Plus in the early days, they were recording the same small sentence pool over and over again, so the first 700 hours or so are duplicates.
I hope there will be efforts in the future to clean up the existing dataset to improve its quality.
But you are right, the process has some flaws. Maybe we can review the dataset automatically on some common errors, once an STT system is ready for a language?
The only other option I can think about is a validation process that includes more people per sentence. Right now, only two people validate a sentence, and if they disagree a third person decides. We could at least double check sentences with one "no" vote one more time.
However, Hillary, the new community manager, seems good and she’s making a lot of positive changes so hopefully this will be addressed soon.
Long-term the best approach may be some kind of user onboarding before they can record / validate.
Thank you for the compliment and feedback.
Following community feedback voice validation criteria is now available on Common Voice platform (released as part of the recent dataset).
This is one of many steps we are making to improve Common Voice contributors and everyone using the dataset.
I'm going to disagree that there's a universal need for perfect training data in ASR. I'm sure it helps with some model types and training processes, but it simply hasn't been a factor in my use of Common Voice (English). I'll also note my best model can hit around 10% WER on Common Voice Test without any language model, which is better than any public numbers I've seen posted for it so far (I'm not even using a separate transformer decoder or RNN decoder layers for this number, just the raw output of CTC greedy decode).
None of the above even factors in techniques like wav2vec and IPL (iterative pseudo labeling) with noisy student, which suggest you can hit extremely competitive accuracy with very little correctly labeled data. These techniques are the underpinnings of the current state of the art models.
It's also impossible (?) to undo a clip. Eg.: If I've already recorded 3 clips and mistakenly begin a clip I simply can't pronounce correctly, there's no way of removing that clip without discarding the whole set. (EDIT: it is possible by re-recording that clip and pressing skip)
Would love to hear from experienced practitioners and a bit of detail on the experience.
Thanks HN community!
Mozilla Deep Speech is an open source speech recognition engine, based upon Baidu's Deep Speech research paper[2].
Unsurprisingly, Deep Speech requires a corpus such as... Common Voice.
It also looks like Baidu are now developing their Deep Speech as open source? https://github.com/PaddlePaddle/DeepSpeech
We could possibly give the developer the benefit of the doubt that they're not doing anything inappropriate with the data but frankly why pass your data through a third party that's not part of the project.
And why install an app requiring access to your shared local storage? The GitHub repo claims the website an animations are slow which sounds like BS to me. It works fine on a five year old phone I use for submitting.
Just contribute here if you're so inclined, much more sensible:
And your point makes little sense, because if the site was not working how could the app get voice data into the project. I've had some involvement with these projects over the years so I'm not just firing off arm-chair comments on this. They wouldn't have been able to add this new voice data if the site was under developed as you imply.
Sure, you can get the source but as I said it's still a pointless step to go via a third party
If so, we would like to use it for the HEAR NeurIPS competition: https://github.com/microsoft/DNS-Challenge/tree/master/datas...
The challenge is restricted only to classification tasks, and sequence modeling like full ASR is unfortunately beyond the scope of the competition.
I have tried pretty much every API offered by big tech, and also various open source models. All of them seem to have incredibly high word error rates. This is mostly for conversations with various Indian accents.
https://alphacephei.com/vosk/models/vosk-model-en-in-0.4.zip
In case you want more accuracy you can share a file with an example, we can take a look on how to make the best accuracy.
For Indian ASR it is also worth to mention recently introduced Vakyansh project which builds model for major Indian languages:
How did they get almost as much training for Kinyarwanda as they have English?
I think it should be explained that one should speak naturally when reading the lines.
news from the past about this :
Initial Release of Mozilla’s Open Source Speech Recognition Model and Voice Data : https://news.ycombinator.com/item?id=15808124
Mozilla releases the largest to-date public domain transcribed voice dataset https://news.ycombinator.com/item?id=19270646
https://commonvoice.mozilla.org/en/datasets
The ratio of male to female tagged voices in the English dataset is 45 percent male to 15 percent female. (The remaining 40 percent is untagged.) Odds are good that the ratio is closer to 75 25 than 50 50, at least by hours of recorded audio.