Open Speech Recording
aiyprojects.withgoogle.com
aiyprojects.withgoogle.com
Why help an ad company when you can help a constant warrior for privacy, freedom, and open source technology? Why is Google trying to create their own instead of joining Mozilla's work?
A HN headline yesterday explained which licenses were prohibited by Google:
https://opensource.google.com/docs/thirdparty/licenses/#bann...
Most of the licenses on the Mozilla page are either restricted or banned at Google, based on a cursory crosscheck.
Text of CC0 for discussion purposes, is here: https://creativecommons.org/publicdomain/zero/1.0/legalcode
CC0 appears to have most of the legal text expected in a proper license including clear allowance of modification and distribution, and lack of warranty or liability.
It's also interesting that the agreement states, "Google may use the clips ... share the clips with others, including the general public, for example, as part of a public dataset to facilitate research" instead of Google will share the data.
Compare that to Mozilla's Common Voice project which has no agreement.
- Common Voice does have a TOS. It's here:
https://voice.mozilla.org/en/terms
- You're equivocating on the word "may". You're reading it as declaration of intent—i.e., as a synonym for the word "might" and def 5 as listed on Wiktionary[1]. In reality, it's being used in line with the most common tendency for things like TOS, RFCs, etc, which is as a synonym for the word "can", i.e., to acknowledge what one is allowed to do (in this case Google), as in def 2.
As to "may", you're incorrect. I'm implying that they should have used stronger language, as in "should" or "will", since otherwise the data contributed to this project may not be made public.
As far as "may" vs "should"/"will", that'd be very unlikely for them to put into the terms: what happens if they decide tomorrow to abandon the project and just leverage a different data set? They'd now be beholden to release something.
Some suggestions for improvements were made in a discussion on the Kaldi Github (https://github.com/kaldi-asr/kaldi/issues/2141), but the bottom line is that it's not clear what the exact use-case of the corpus is. It's not command recognition, there's no particular domain, just more or less random English sentences collected from CC0 sources. Use cases should drive corpus collection.
That's almost same thing you say to vocally authorise yourself with HMRC in the UK...
It's something like "My voice is my password".
I wouldn't want me saying that phrase released into public...
https://news.ycombinator.com/item?id=15808124
https://blog.mozilla.org/blog/2017/11/29/announcing-the-init...