I'm not sure I want the latest twitter trend to be involved in the design of my translator...
I'm not sure I want the latest twitter trend to be involved in the design of my translator...
There are more details in the paper if you want and the mitigation code is all open source if you want to check what it actually does.
Oh, well that clears it up! </snark>
I don't see any definition of 'toxicity' on the landing page - it seems to be one of those 'I know it when I (hear) it' kind of words... unless there's some widely-accepted definition in this area of study?
The tldr is that if you say: "Thank you for this job offer." you wouldn't want it to be (mis)translated as "Go F*k yourself.". But if you do say "Go F yourself", you still want it to be translated as that.
The Hitchhiker's Guide To The Galaxy claims the opposite:
"Meanwhile, the poor Babel fish, by effectively removing all barriers to communication between different races and cultures, has caused more and bloodier wars than anything else in the history of creation."
e.g. "geil" (either cool or horny depending on usage) in German
It's not fundamentally different than e.g. "wicked" in English, but the biggest bias that potentially all these ML models exhibit is predisposition towards Anglophoneism
You can find details on how the multi-language creation of the toxicity lists was done in section 7.3 of the NLLB paper: https://arxiv.org/pdf/2207.04672.pdf. TLDR: it's not just a translation of a base English list, even if we started from that, each language has a curated list that was built by professional translators.
Can it make sure that the output toxicity level is not lower than the input?
If not (which I strongly suspect is the case), then that is unacceptable. We cannot fight toxic narratives with ignorance.
There is no moral superiority to deny or force label other people's identities. You're an attack helicopter? Great, roger dodger, let's go get coffee Seahawk.
No one is seriously asking for litter boxes in school bathrooms or helicopter refueling stations.
This feels a bit out-of-nowhere.
My read on parent comment was that "Twitter trends" are fast-changing norms about what language is (un)acceptable. They were not saying that LGBTQIA+ identity itself is a trend.
The original comment you replied to made the point that they don't want their own personal expression curtailed or modified according to someone else's opinion of acceptable speech.
As someone who repudiates Russia's policies, I support and agree with their point.
Taking what they wrote as harshly as possible, a translation model's output might include narrative elements from the transphobic judgement you are concerned about. That would be a problem, because it would amplify transphobic narratives.
Taking what they wrote as favorably as possible, a translation model's output might rephrase what was written, such that a pro-LGBTQIA+ inclusion narrative is more eloquently expressed than the author actually intended. That would be a problem, because hiding the reality of transphobic narratives would remove our ability to recognize and talk about them.
To make this even more complicated, what if we are using this model for real-time dialogue? What happens when someone says something vaguely transphobic, their words get translated to an inclusive narrative, and you continue that inclusive narrative in your reply? Should the translator alter your words to be transphobic? If it doesn't, then will the entire conversation go off the rails, or will both parties continue, oblivious of each others' ideological subtleties?
---
I don't believe for a second that a model could be trained to avoid toxic narrative and translate accurately.
Hallucination is a feature, not a limitation. The sooner "AI" narratives can accept this reality, the better.
Perhaps it's taking a Big List of Naughty Words and weighting them so that the system must be "extra sure" that's what the speaker said, or else fall back to a G-rated word?
1: https://www.google.com/search?q=engrish+fucking+sign&tbm=isc...
From the hackernews guidelines