Maybe you could convince a couple of indie creators or state-run programs to licence their audio? But I'm not sure if negotiating that is more efficient than just recording a bit more audio, or promoting the project to get more volunteers.
Also I don't think there have been any major court cases about this, so there's no clear precedent in either direction.
Easy fix - keep a bloom filter of hashed ngrams ensuring you don't repeat more than N words from the training set.
It would still be an uphill battle to convince them to hand over the training set but the legal department can likely be convinced if the data set they contribute back is heavily chopped up audio of the original content, especially if they have the originals before mixing. I imagine short audio files without any of the music, sound effects, or visual content are pretty much worthless as far as IP goes.
I'm not sure that is a clear copyright violation. Sure, at a glance it seems like a derivative work, but it may be altered enough that it is not. I believe that collages, and reference guides like cliff notes are both legal.
I think a bigger problem would be that the scripts, and even the closed captioning, rarely match the recorded audio 100%
What might work better is using closed captions or subtitles, but I've also seen enough cases where those don't exactly match the actual speech either.
Thanks for taking time to contribute!
(The first clip I got was spoken more or less correctly, but a couple of words are slurred together and the prosody is awkward. Without having a good idea of the standards and goals of the project, I have no idea whether including this clip would make the overall dataset better or worse. My gut feeling is that it's good for training recognition, and bad for training synthesis.)
This seems to me like a major issue, since it should take a relatively small amount of effort to write up a list of guidelines, and it would be hugely beneficial to establish those guidelines before asking a lot of volunteers to donate their time. I don't find it encouraging that this has been an open issue for four years, with apparently no action except a bunch of bikeshedding: https://github.com/common-voice/common-voice/issues/273
[1] https://f-droid.org/packages/org.commonvoice.saverio/
[2] https://discourse.mozilla.org/t/discussion-of-new-guidelines...
One speaker, who sounded like they were from the mid-west United States, was dropping the S off words in a couple clips. I wasn't sure if it was misreads or some accent I'd never heard.
Another speaker, with a thick accent that sounded European, sounded out all the vowels in circuit. Had I not had the line being read, I don't think I'd have understood the word.
I heard a speaker with an Indian accent who added a preposition to the sentence that was inconsequential but incorrect none the less.
I hear these random prepositions added as flourishes frequently with some Indian coworkers, does anyone know the a reason? It's kind of like how American's interject "Umm..." or drop prepositions (e.g. "Are you done your meal?") and I almost didn't pick up on it. For that matter where did the American habit of dropping prepositions come from? It seems like it's people in the North East primarily.
[¹] If that's even fair given it's a dialect in its own right - Americans also say things differently than I would as a 'Britisher'
I'd say it's mostly subtler (I suppose that should be the expected distribution!) things I've noticed though, they're just harder to recall as a result.
(Just want to emphasise I'm not making fun of anybody or saying anything's wrong, in case it's not clear in text. I'm just enjoying learning Hindi, fairly interested in language generally, and interested/amused to notice these things.)
Hindi is much more economical, to put it literally, one says things like 'than/from/compared to orange, lemon is sour', and 'orange is little/less [without comparison] sour'.
Which, I believe, is what gives rise to InE sentences like 'the salt in this is very less' (it needs more salt, there's very little).
Massive mountains of data tends to be incompatible with opensource projects. Even Mozilla collecting user statistics is pretty controversial. Imagine someone like Mozilla trying to collect hundreds of voice clips from each of tens of millions of users!!
Both of those involve entering data about external things. Asking people to share their own data is another thing entirely—I suspect most people, me included, are much more suspicious about that.
Do user collected clips have soemthing so special to the point that it’s critical to collect them?
They do, and it's working! https://commonvoice.mozilla.org/en
It's well-documented and works basically out of box. I wish the STT models bundled were closer to the quality of Kaldi but the ease-of-use has no comparisons.
And maybe with time it will surpass Kaldi in quality too.
CMUSphinx and Julius have been around for ~10+ years at this point.
[EDIT] - there's even a useful Quora post[4]
[0]: https://cmusphinx.github.io/
[1]: https://www.kaldi-asr.org/doc/about.html
[2]: https://github.com/alphacep/vosk-server
[3]: https://github.com/julius-speech/julius
[4]: https://www.quora.com/Are-there-any-open-source-APIs-for-spe...
Decent designs are in published papers all over the place, so thats a solved issue.
Lots of compute requires lots of $$$, which isn't opensource-friendly.
Lots of data also isn't really opensource friendly.
Sadly this is a niche that the opensource business model doesn't really fit.
* besides the hard part of standing up an EMR is not installing a prepackaged software.
Not really, look up BOINC.
A FLOSS system would only have my voice to recognise and I would be willing to spend some time training it. Very different usecase from a massive cloud that should recognise everyone's voice and accent.
Example: https://www.youtube.com/watch?v=tfcme7maygw
Granted, maybe this is "not good enough", but I feel like I got pretty far with pico2wave, pocketsphinx plus 1980's Zork level "comprehension" technology.
And the open source status of pico2wave is a bit questionable, I'll grant you that.
A bit more detail about the implementation here: https://scaryreasoner.wordpress.com/2016/05/14/speech-recogn...
So companies with a lot of data, then.