First AI model that translates 100 languages without relying on English
about.fb.com
about.fb.com
We worked with the college press office to draft our press release, and in it we made the claim that it was the first wind turbine on a school campus in the greater Boston area.
Well one of the bigger papers picked up the story and the press office ended up getting a nasty voicemail from some parent at a private high school in some Boston suburb. The parent wanted our school to know that HIS kid's school was the first to put up a wind turbine.
This actually worried our little team because we didn't want to say something untrue. But the pr person explained that this "first to do X" type of thing is a very common press tactic and it is not unusual to have your claim disputed.
And that second, "greater Boston area" is not really a narrowly defined term.
This was something of a relief and we sort of had a running gag about the PTA dad with sour grapes about our wind turbine from then on.
You can call it hyperbole and good marketing, but then it also goes to the extent of PR after oil spills, tailings damns wiping out villages and poisoning or risking the health of your own workforce, like for example Dupont and C8.
It's hard not to be cynical about PR because you cannot do it effectively without creating / propagating a distorted view of the truth.
But then again, the reason it still essentially needs to exist is because there is no universal truth, no universal sense of good or bad. So for example, do I forgive nuclear power and GM food companies from trying to influence public discourse on the safety of their work? Of course, it's essential in a world where they are competing against lots of unwarranted negative attention, much of it baseless as far as the science goes.
Anyway in spite of my ranting and bias, really I'd conclude it's a necessary evil, and the majority of the industry is innocuous and just trying to showcase the work and values of their clients...
Also newspaper editors are a huge part to blame, when these days loads of press releases make it to print almost unedited (and as has happened to myself and others I know, the bits they changed were even more wrong - including in national level newspaper)... Fact checking claims and sources is so lacking, especially in fast modern news cycle.
[edit] typo
That may be true to some extent if you really push the philosophy angle, but I'd say real world is much simpler than that. There's stuff that happened, and there are consequences - just because the consequences may not be fully computable in the amount of time and effort anyone is willing to expend on it, doesn't mean truth suddenly becomes fuzzy. The territory is sharp, it's just the map that's uncertain. But for that, the proper words are "I'm not sure", not "I have my truth, you have yours".
> So for example, do I forgive nuclear power and GM food companies from trying to influence public discourse on the safety of their work? Of course, it's essential in a world where they are competing against lots of unwarranted negative attention, much of it baseless as far as the science goes.
And because of what I wrote above, I despise both. Yes, I understand the practical necessity - one side lies because the other side lies too, both stuck in a feedback loop. But I'd still say both are behaving unethically. Two wrongs don't make a right.
I subscribe to the viewpoint I've best seen phrased in an old blog post[0]: "promoting less than maximally accurate beliefs is an act of sabotage. Don’t do it to anyone unless you’d also slash their tires".
--
[0] - http://web.archive.org/web/20080915221100/http://www.acceler...
Oh but it definitely DOES mean that truth becomes fuzzy. There is NO territory we can meaningfully talk about outside our subjective maps. Likewise, there is no such thing as "absolute truth" - and I do mean this on a very practical, day-to-day level, not in an abstract philosophical way.
The sooner we accept and embrace this, the sooner we can move away from "I'm right and you're wrong" to "let's make progress on what actually matters to the parties involved".
I wonder if part of the problem is people do not have time or cultural persuasion to enjoy philosophical consideration of matters as they do tawdry headlines.
In this I agree. Looking back on it now, if I some other school had already put in any sort of turbine, I would have looked for another headline.
At the time, our group wasn't aware of this other school's efforts. It felt like we were all over cleantech in New England, so the dispute was a genuine surprise.
I pay close attention to the misuse of facts or lack of context around facts now. I try to look for the truth even if it isn't the way I might want it to be.
When I find out the truth is counter to the editor's headline on news, I get upset enough to try and pinpoint where the spin is coming from.
Granted, this was only one afternoon at an event organized by the Dutch Judo association.
Depending on the circumstances, I omit the latter statement.
This is spin.
Will not happen in pure AI-based system without hard-coded parsing rules and without changing the obviously default dictionary-based english-like language model.
Maybe some texts that describe a given language should also be used.
This doesn't seem an insurmountable problem.
Basically a full training set with all the words does not exist. Even less a set with actual translations for them.
This also means that autocomplete etc is pretty much useless for us. I just disable it on iPhone as it actively makes typing out a sentence harder.
(also to make things harder as others have pointed out the base word and inflections get changed to make the word easier to pronounce so there is no static form for them when you stack them)
For example here is some (but not all) for the word dog (Koira). The longer ones would be full sentences on their own in most languages.
Koira, koiran, koiraa, koiran again, koirassa, koirasta, koiraan, koiralla, koiralta, koiralle, koirana, koiraksi, koiratta, koirineen, koirin, koirasi, koirani, koiransa, koiramme, koiranne, koiraani, koiraasi, koiraansa, koiraamme, koiraanne, koirassani, koirassasi, koirassansa, koirassamme, koirassanne, koirastani, koirastasi, koirastansa, koirastamme, koirastanne, koirallani, koirallasi, koirallansa, koirallamme, koirallanne, koiranani, koiranasi, koiranansa, koiranamme, koirananne, koirakseni, koiraksesi, koiraksensa, koiraksemme, koiraksenne, koirattani, koirattasi, koirattansa, koirattamme, koirattanne, koirineni, koirinesi, koirinensa, koirinemme, koirinenne, koirakaan, koirankaan, koiraakaan, koirassakaan, koirastakaan, koiraankaan, koirallakaan, koiraltakaan, koirallekaan, koiranakaan, koiraksikaan, koirattakaan, koirineenkaan, koirinkaan, koirako, koiranko, koiraako, koirassako, koirastako, koiraanko, koirallako, koiraltako, koiralleko, koiranako, koiraksiko, koirattako, koirineenko, koirinko, koirasikaan, koiranikaan, koiransakaan, koirammekaan, koirannekaan, koiraanikaan, koiraasikaan, koiraansakaan, koiraammekaan, koiraannekaan, koirassanikaan, koirassasikaan, koirassansakaan, koirassammekaan, koirassannekaan, koirastanikaan, koirastasikaan, koirastansakaan, koirastammekaan, koirastannekaan, koirallanikaan, koirallasikaan, koirallansakaan, koirallammekaan, koirallannekaan, koirananikaan, koiranasikaan, koiranansakaan, koiranammekaan, koiranannekaan, koiraksenikaan, koiraksesikaan, koiraksensakaan, koiraksemmekaan, koiraksennekaan, koirattanikaan, koirattasikaan, koirattansakaan, koirattammekaan, koirattannekaan, koirinenikaan, koirinesikaan, koirinensakaan, koirinemmekaan, koirinennekaan, koirasiko, koiraniko, koiransako, koirammeko, koiranneko, koiraaniko, koiraasiko, koiraansako, koiraammeko, koiraanneko, koirassaniko, koirassasiko, koirassansako, koirassammeko, koirassanneko, koirastaniko, koirastasiko, koirastansako, koirastammeko, koirastanneko, koirallaniko, koirallasiko, koirallansako, koirallammeko, koirallanneko, koirananiko, koiranasiko, koiranansako, koiranammeko, koirananneko, koirakseniko, koiraksesiko, koiraksensako, koiraksemmeko, koiraksenneko, koirattaniko, koirattasiko, koirattansako, koirattammeko, koirattanneko, koirineniko, koirinesiko, koirinensako, koirinemmeko, koirinenneko, koirasikaanko, koiranikaanko, koiransakaanko, koirammekaanko, koirannekaanko, koiraanikaanko, koiraasikaanko, koiraansakaanko, koiraammekaanko, koiraannekaanko, koirassanikaanko, koirassasikaanko, koirassansakaanko, koirassammekaanko, koirassannekaanko, koirastanikaanko, koirastasikaanko, koirastansakaanko, koirastammekaanko, koirastannekaanko, koirallanikaanko, koirallasikaanko, koirallansakaanko, koirallammekaanko, koirallannekaanko, koirananikaanko, koiranasikaanko, koiranansakaanko, koiranammekaanko, koiranannekaanko, koiraksenikaanko, koiraksesikaanko, koiraksensakaanko, koiraksemmekaanko, koiraksennekaanko, koirattanikaanko, koirattasikaanko, koirattansakaanko, koirattammekaanko, koirattannekaanko, koirinenikaanko, koirinesikaanko, koirinensakaanko, koirinemmekaanko, koirinennekaanko, koirasikokaan, koiranikokaan, koiransakokaan, koirammekokaan, koirannekokaan, koiraanikokaan, koiraasikokaan, koiraansakokaan, koiraammekokaan, koiraannekokaan, koirassanikokaan, koirassasikokaan, koirassansakokaan, koirassammekokaan, koirassannekokaan, koirastanikokaan, koirastasikokaan, koirastansakokaan, koirastammekokaan, koirastannekokaan, koirallanikokaan, koirallasikokaan, koirallansakokaan, koirallammekokaan, koirallannekokaan, koirananikokaan, koiranasikokaan, koiranansakokaan, koiranammekokaan, koiranannekokaan, koiraksenikokaan, koiraksesikokaan, koiraksensakokaan, koiraksemmekokaan, koiraksennekokaan, koirattanikokaan, koirattasikokaan, koirattansakokaan, koirattammekokaan, koirattannekokaan, koirinenikokaan, koirinesikokaan, koirinensakokaan, koirinemmekokaan, koirinennekokaan
Or the shop (Kauppa) one linked in the thread already http://www.ling.helsinki.fi/~fkarlsso/genkau2.html
We inflect pretty much all words (verbs, nouns, pronouns, numerals, adjectives and some particles)
https://forum.ultras-tifo.net/countries-ball-39-s-comics-t28...
That doesn't mean it's impossible to work with other languages, just that "words separated by spaces" is the wrong abstraction for processing them. It just happens to be a heuristic that works well enough for English, so a lot of functionality (like autocomplete) assumes that it works the same for other languages. It would be perfectly feasible to offer partial completions of long words in languages like Finnish or German, if only the space key were treated as less special. (Just compare to Chinese and Japanese, where autocomplete works despite no spaces at all.)
Not having a static form might create some redundancy in the lexicon, but that's not more of a problem than the vowel mutation in English "sing", "sang", "sung", "song". Treating different surface realizations of the same underlying base form as independent might actually be beneficial for getting accurate results that take into account how the base form is modified.
But in what language would those texts be? Language models used for translation are trained on sentence pairs. How would e.g. a book on the grammar of Finnish, written in Finnish and without a translation in English, help learn how to translate Finnish into English?
I'm genuinely asking. This sounds like an interesting idea- but how exactly would it work?
https://en.wikipedia.org/wiki/Georgian_grammar
Problem is of course how to learn the meanings, when they are omitted in translations. Even native speaker is not always sure, like is "koiran-ko-han-ko" a tautology, or does the second question-mark refer to to the "han". "You wanted a dog, yes?" vs. "You wanted a dog, yes? You sure now?".
I'm almost sure autocompletion in Finnish (or Estonian or Hungarian) is borderline useless, as the chances to write a word that wasn't ever written before are quite high.
But even with more "sane" grammars these models barely work. When I write in Spanish, often there are verb forms that are missing that I have to type fully.
For example, in Spanish, all forms are on conjugation lists, but it can be tough to find a set of corpuses that covers all. I know that a 2016 Spanish Wikipedia dump I played with covered about 20% of all verb forms present in rae.es.
Then on all those forms (I'd say about 16-18 tense/mood/aspect forms are in common use) you have to take into account enclitic/affix pronouns, and the whole thing goes awry quick.
"Comámosnoslas" - If you search on Google, there are only 9 results[0], this post likely to become the 10th. But it's a completely normal word a Spanish speaker may use and will understand.
comamos ("let us eat") nos (emphasis "for ourselves") las ("them", feminine). "Let's eat them!" but with emphasis lost in translation.
E.g. usage "Hay 3 pizzas, ¡comámosnoslas!" . This sentence, funnily, is properly translated by Google Translate into English, but it tries to correct it to "comamosnos las" which is absolutely broken Spanish.
I agree with you in that regard, speaking Turkish.
In terms of life, generally in Helsinki or maybe the other big cities you could get by just fine with english only. I know a few folks at work who've been here a long time and have survived ok! In terms of written information everything is in Finnish or Swedish but occasionally (and from what I gather, increasingly) in English. I've found pretty much everyone in a customer/client facing position speaks very good English too.
The one situation that comes to mind where it's genuinely quite difficult is grocery shopping and following cooking instructions. Since I like to cook I enjoy trying new things but Google translate is totally hopeless. Thankfully I've found translating the Swedish to English seems to work OK.
They are both neighbouring contries of the UK, with similar idioom. Finnish is another league. Congrats BTW, many English native speakers these days try but, in my experience, use English in business settings only.
According to this metric, if you have a moderately long sentence like "I am not the person who said the president should be reelected" and your translation missed the "not", you would still get a score of 11/12 ~ 92%. And, as far as I know, word order doesn't even matter, so "I am the person who said the president should not be reelected", while wrong, would get a perfect score.
Of course these are rather artificial examples, and in general machine translation algorithms and their evaluation work because it's "easier" to create an algorithm that gets the right translation than one that, unintentionally, fools the metric systematically. Nevertheless if the research community used a metric that punished this kind of mistakes more strongly, I suspect that over time a few new algorithms could come up that improve on this specific point.
Alas, I don't know of any such metric (nor I would know how to design one, of course, otherwise I'd publish it ;-) ).
It is also just false that BLEU does not care about order.
It's an extremely common idiom and I used it in a proper context. It fails, horribly. I don't know why I even bother checking on Google Translate. It basically fails every single time to create natural Japanese.
Our teachers could tell in a second if it was made by Google Translate.
Just don't rely on it. Ever.
Tamil isn’t an Indo-Aryan language. It would be rather surprising if it were a good bridge for any two Indo-Aryan langauges.
https://dl.fbaipublicfiles.com/m2m_100/12b_last_checkpoint.p...
I can't imagine the amount of computation it took to generate this.
They have existing tools to jointly embed sentences from multiple languages (laser). These models are trained using parallel corpora involving either english or spanish on a translation task.
Using these models and the joint embeddings they produce, fb can mine the web for new pairs by roughly identifying whether two sentences in two different languages correspond to a translation pair.
Still, it lets you massively increase the amount of data you're training on, which is generally worth it.
A more ELI5 explanation would be something like this: Laser is an encoder/decoder architecture trained as a translation task from language X to english/spanish, with the particularity of having only one vector between the encoder and decoder. Once this system is trained (with public translation datasets), the vector that the encoder gives you for an input sentence represents the "meaning" of that sentence, since the decoder relies only on that information to generate a translation. And this vector representation is the same for any language the system was trained on.
So we use that system to mine data from commoncrawl: giving any language pair (say romanian-nepali), having two vectors close to each other in that latent space means the sentences have the same meaning.
We use fastText's language classifier [1] to filter from commoncrawl, compute the vector representations with Laser, and find close vectors thanks to Faiss [2].
[0] https://engineering.fb.com/ai-research/laser-multilingual-se...
[1] https://fasttext.cc/docs/en/language-identification.html
That's the goal, but have you asked any Romanian-Nepali bilinguals how well it works? (I realize those might be hard to come by.) I had a look at some of the language pairs in CCMatrix, and I noticed that the highest-confidence matches for some of them (English-Chinese, English-Korean) include a lot of quotations from old religious texts. That wouldn't be a problem if they were actually the same quotations, but it looked more like the model managed to identify archaic language and then got overly confident that two archaic sentences must have the same meaning.
I wonder whether there's been a human evaluation of the mined training data, or whether you rely on catching any problems downstream when you measure the BLEU of the trained model.
http://www.cervantesvirtual.com/obra-visor/el-ingenioso-hida...
Edit: to clarify, other translation servies I've used, notably Google, do translate from and to Greek but make a meal out of it.
Douglas Hofstadter's test sentence: In their house, everything comes in pairs. There’s his car and her car, his towels and her towels, and his library and hers.
DeepL translation: Dans leur maison, tout vient par deux. Il y a sa voiture et la sienne, ses serviettes et les siennes, sa bibliothèque et la sienne.
DH translation: Chez eux, ils ont tout en double. Il y a sa voiture à elle et sa voiture à lui, ses serviettes à elle et ses serviettes à lui, sa bibliothèque à elle et sa bibliothèque à lui.
As a French native, I can tell you that DH's translation is great. And DeepL's reply is still something you can recognize at automatic translation at first sight.
Don’t get me wrong, more often than not, current MT tools give good enough results to understand the treated topic. But that’s still hideous translations. I wouldn’t buy a book translated through them for example.
Not that I'd expect the computer to "know" that. I'm just curious, as a French learner, about why DH's translation is better.
What do you mean with "a reason why"? From a synchronic or a diachronic point of view? Yes, there are reasons that linguists can provide. But for the mere layman, it will simply be weird to encounter such an utterance, full stop. Not that your question is non-sense, I would rather simply expect most people to be unable to tell you why despite they know it sounds weird.
> I know it's just a word-for-word translation, but the latter seems more natural, and the former unnecessarily verbose.
What you mean with "seems more natural", is that it sounds closer to what you are accustomed to in your own language. French is generally more verbose than English, especially in usual written form.
Now, as your "why" is nontheless completely legitimate, here is one explanation .In the case of "sa bibliothèque et la sienne", unlike in English which insist on the possessor (his/her takes the gender of the one who owns the object), French insists on the possessed (sien/sienne takes the gender of the object which is possessed, and every substantive have a gender). So would you translate "her library and his" in the same way, you would end up with the very same translation "sa bibliothèque et la sienne". On the other hand "sa bibliothèque à elle et sa bibliothèque à lui" would preserve the information of who owns each library.
Would you really mind to comes that is as close as possible to the original prosody, you might use "sa bibliothèque à lui, et elle la sienne."
Text to speech is also much worse in deployment compared to research. The recent research models have much better intonation.
GPT-3 is the worst offender here - it's so big that it becomes almost uneconomical to run, and certainly impossible to offer for free. (estimated requirements are 11 Tesla V100 GPUs)
Afrikaans Albanian Amharic Arabic Armenian Asturian Azerbaijani Bashkir Basque Belarusian Bengali Bosnian Breton Bulgarian Burmese Catalan Cebuano Chinese Croatian Czech Danish Dutch Eastern Punjabi English Estonian Finnish French Fulah Galician Ganda Georgian German Greek Gujarati Haitian Hausa Hebrew Hindi Hungarian Icelandic Igbo Ilokano Indonesian Irish Italian Japanese Javanese Kannada Kazakh Khmer Korean Lao Latvian Lingala Lithuanian Luxembourgish Macedonian Malagasy Malay Malayalam Marathi Mongolian Nepali Northern Sotho Norwegian (Bokmål) Occitan Oriya Oromo Pashto Persian Polish Portuguese Romanian Russian Scottish Gaelic Serbian Sindhi Sinhalese Slovak Slovenian Somali Spanish Sundanese Swahili Swati Swedish Tagalog Tamil Thai Tswana Turkish Ukrainian Urdu Uzbek Vietnamese Welsh West Frisian Wolof Xhosa Yiddish Yoruba Zulu
Some pairs are also very bad right now, particularly Korean->English.
IMO, this is a dataset problem. I think we're finally seeing the "next generation" of dataset collection in AI so there will be hopefully improvements for pairs with sparser sets.
DeepL makes perfect translations between English and German. Perfect in a sense that it looks like a professional translator translated it. Google, Microsoft, Facebook might be far off, but deepl isn't.
The possibility to insert your corrections in the suggested translation and let the algorithm rewrite the following text make it a huge time saver when one needs to do a translation.
1. https://en.wikipedia.org/wiki/Google_Neural_Machine_Translat...
Otherwise maybe https://huggingface.co/ or similar might release a version with interface.
One day it will tell us what Mark Z is really saying.