Reminder that Mozilla's Common Voice project accepts voice donations! https://voice.mozilla.org/
Reminder that Mozilla's Common Voice project accepts voice donations! https://voice.mozilla.org/
Kabylia is a region in the north of Algeria mostly inhabited by Berber people who are bilingual in Algerian Arabic and Kabyle. In recent years, an independence movement has developed that emphasizes Kabyle over Arabic for reasons of internal cohesion. To confuse matters, there's also a pan-Berber movement denying the existence of a separate Kabyle language, classifying it as a dialect of Berber/Tamazight instead.
Those heated politics have led to a large number of Kabyles contributing to various linguistic corpus projects to gain visibility for their cause. E.g. trying to overtake Berber on https://tatoeba.org/stats/sentences_by_language (As far as I know, Mozilla's Common Voice shares data with the Tatoeba project.)
https://en.wikipedia.org/wiki/LibriVox
The recordings are public domain audio books of public domain books, so the licensing should be fine. The audio isn't annotated, but given the value involved I think it would be worth attempting to use forced alignment to annotate the recordings with their public domain source texts. Forced alignment using the sort of speech recognizer you're trying to train in the first place may be a bit "chicken and the egg", but from some experiments I've run myself existing open source speech recognizers can do it reasonably well. Humans could manually tune up the alignment to improve the quality if necessary.
As for motivating people to actually do that mundane work... well these are audio books so maybe the work isn't so mundane after all! The LibriVox recording of Tom Sawyer (read by John Greenman: https://librivox.org/tom-sawyer-by-mark-twain/) is pretty great and has been listened to by millions of people. If somebody created a "read along" web app that showed you the text of the book from Project Gutenberg getting highlighted as the audiobook from LibriVox was played, users who have an interest in reading/hearing the book could have their attention held by Mark Twain and with the right UI provide fine tuning for the forced alignment at the same time.
[1] Panayotov, V., Chen, G., Povey, D. and Khudanpur, S. (2015). LibriSpeech: an ASR corpus based on public domain audio books. Proc. ICASSP. http://www.danielpovey.com/files/2015_icassp_librispeech.pdf