5 BLEU points more is massive with 4 times less params.
The fact that layers themselves are narrower means that training and evaluation of the NN is also much faster.
The fact that layers themselves are narrower means that training and evaluation of the NN is also much faster.
Since the network will have seen far more english text than text of other languages, it suggests that performance on limited training data is more improved.