HNHacker News
TopNewBestAskShowJobs

hello_im_angela

71 karma · joined July 6, 2022

submissionscomments
hello_im_angela··on No Language Left Behind
We represent all languages in their natural script, rather than transliterating them into a common synthetic one.

Regarding Mari: extremely interesting language, exciting to hear that you are from that region. We are interested in working on this one (likely in the "Hill Mari" variant), but currently do not support it.

hello_im_angela··on No Language Left Behind
We tokenize with the flores-200 spm model, correct. To generate from the model, check out the instructions here: https://github.com/facebookresearch/fairseq/tree/nllb/exampl...
hello_im_angela··on No Language Left Behind
could you send me an email please? It's available on our paper, page 1: https://research.facebook.com/publications/no-language-left-...

Regarding grants: we have offered compute grants previously with the Workshop on Machine Translation (last year: https://www.statmt.org/wmt21/flores-compute-grants.html, this year: https://statmt.org/wmt22/large-scale-multilingual-translatio...) and we have an RFP, but it's currently focused on African languages: https://ai.facebook.com/research/request-for-proposals/trans...

hello_im_angela··on No Language Left Behind
sooo real. Many low-resource languages have many different natural variants, can be written in multiple scripts, don't have as much written standardization, or are mainly oral. As part of the creation of our benchmark, FLORES-200, we tried to support languages in multiple scripts (if they are naturally written like that) and explored translating regional variants (such as Moroccan Arabic, not just Arabic).

As an aside, the question of how to think about language standardization is really complex. We wrote some thoughts in Appendix A of our paper: https://research.facebook.com/publications/no-language-left-...

hello_im_angela··on No Language Left Behind
ha yes, that's correct. If you have thoughts on specific constructed languages where having translation would really help people, let us know!
hello_im_angela··on No Language Left Behind
We interviewed speakers of low-resource languages from all over the world to understand the human need for this kind of technology --- what do people actually want, how would they use it, and what's the quality they would find useful? Many low-resource languages lack data online, but are spoken by millions. However, many indigenous languages are spoken by smaller numbers of people, and we are definitely interested in partnering with local communities to co-develop technology and have been actively investigating these collaborations but don't have much to share yet.
hello_im_angela··on No Language Left Behind
We release several smaller models as well: https://github.com/facebookresearch/fairseq/tree/nllb/exampl... that are 1.3B and 615M parameters. These are usable on smaller GPUs. To create these smaller models but retain good performance, we use knowledge distillation. If you're curious to learn more, we describe the process and results in Section 8.6 of our paper: https://research.facebook.com/publications/no-language-left-...
hello_im_angela··on No Language Left Behind
We have a full list here (copy pastable): https://github.com/facebookresearch/flores/tree/main/flores2... and Table 1 of our paper (https://research.facebook.com/publications/no-language-left-...) has a complete list as well.
hello_im_angela··on No Language Left Behind
It's an extremely difficult problem indeed. A lot of people on the team speak low-resource languages too (my native language as well!), so definitely resonate with what you're saying. My overall feeling is: yeah it's hard, and after decades we can't even do German translation perfectly. But if we don't work on it, it's not gonna happen. I really hope that people who are excited about technology for more languages can use what we've open sourced.
hello_im_angela··on No Language Left Behind
If you're curious to try the system yourself, it's actually being used to help Wikipedia editors write articles for low-resource language Wikipedias: https://twitter.com/Wikimedia/status/1544699850960281601