There are more details in the paper if you want and the mitigation code is all open source if you want to check what it actually does.
There are more details in the paper if you want and the mitigation code is all open source if you want to check what it actually does.
The Hitchhiker's Guide To The Galaxy claims the opposite:
"Meanwhile, the poor Babel fish, by effectively removing all barriers to communication between different races and cultures, has caused more and bloodier wars than anything else in the history of creation."
e.g. "geil" (either cool or horny depending on usage) in German
It's not fundamentally different than e.g. "wicked" in English, but the biggest bias that potentially all these ML models exhibit is predisposition towards Anglophoneism
You can find details on how the multi-language creation of the toxicity lists was done in section 7.3 of the NLLB paper: https://arxiv.org/pdf/2207.04672.pdf. TLDR: it's not just a translation of a base English list, even if we started from that, each language has a curated list that was built by professional translators.
Can it make sure that the output toxicity level is not lower than the input?
If not (which I strongly suspect is the case), then that is unacceptable. We cannot fight toxic narratives with ignorance.
Oh, well that clears it up! </snark>
I don't see any definition of 'toxicity' on the landing page - it seems to be one of those 'I know it when I (hear) it' kind of words... unless there's some widely-accepted definition in this area of study?
The tldr is that if you say: "Thank you for this job offer." you wouldn't want it to be (mis)translated as "Go F*k yourself.". But if you do say "Go F yourself", you still want it to be translated as that.