> It is pretty out of date (ML is a fast moving field) - I doubt Google is using LSTMs for translation in 2020.
Sure, but that doesn't change the fact the training data is focused on English-to-X and X-to-English corpora. The underlying architecture of the model is just an implementation detail that doesn't really affect this as demonstrated by my example.
> These models are trained on Russian-German corpora, but because there is little resources for that, they are supplementing with Russian-English and English-German.
This is exactly what I'd argue isn't the case at all. Otherwise words that have a direct 1:1 translation wouldn't be mistranslated and companies like Yandex wouldn't be able to deliver so much better results.
German-Russian isn't low-resource at all, given 95M and 150M native speakers respectively and a close history for the past 150 years. [edit]The rich cultural history of both countries resulting in a vast library of literature, theatre plays, news publications, films and the general cultural relevance of both languages is even more important.[/edit] It's simply (quite comprehensible) bias towards English for research taking place in the USA and the fact that it's much easier to compile English-to-X and X-to-English corpora in a predominantly English-speaking country.
There are tons of translated books, films, news paper articles, scientific papers, etc. available for Russian-German and Yandex, being a Russian company, naturally has no problem compiling a Russian-German corpus (since they're not biased towards English).