Facebook's AI Just Set a New Record in Translation
forbes.com
forbes.com
https://code.fb.com/ai-research/unsupervised-machine-transla...
There's (edit: what appears to be) an active exploit in their ad network, one that's getting around Chrome's redirect blocking through an apparent 0day.
I'm on Chrome Beta 69.0.3497.53 on Android, so this may not apply outside that.
Chrome team: https://bugs.chromium.org/p/chromium/issues/detail?id=879938
Forbes serves a large portion of their ads in same origin iframes and so is not fully covered by this protection.
[1] https://blog.chromium.org/2017/11/expanding-user-protections...
Can anyone add context to this? Can't seem to wrap my head around this part. Doesn't "he" as a part of a word translate differently in different words?
Perhaps a more appropriate example (just something I thought of, idk if it would work exactly like this) would be like "antiviral" going to "an" "ti" "vi" "r" "a" "l". It could connect "an" and "ti" and associate it with the meaning of "anti" as a prefix along with connecting "vi" and "r" and giving a possible association of "virus". Finally, it could combine "a" and "l" into the suffix "al" and recognize that meaning.
Again, just my two cents. Not particularly sure if this is how that works.
they tokenize text and then learn an embedding for those tokens using (in their case) both the source and target lanuage. This presumably captures regularities for languages that are not very different and has the benefit of having a small dictionary.
For example, polymorphism could be decomposed into poly-morph-ism. Antidisestablishmentarianism, which is unlikely to appear much in the corpus, becomes anti-dis-establish-ment-arian-ism. Now the system can learn how to reuse "anti-" or "establish" from other examples more easily than trying to learn the full word's meaning from the one or two examples it might see in the corpus.
BPE is a clever way to induce these sort of decompositions automatically without any linguistic annotation, making them useful in multilingual settings. Other languages are much more morphologically rich than English, and there it really benefits.
[1] https://www.microsoft.com/en-us/research/wp-content/uploads/...
That sounds almost to good to be true. Excited to see what gets developed with these techniques!
GP's point is still a good one though: while Urdu and English have diverged quite a bit despite being of the same stock, they probably still share a lot more typologically than, say, English and Mandarin or Warlpiri.