BLEU Score: Bilingual Evaluation Understudy
leimao.github.io
leimao.github.io
I tried submitting a NLP paper where I explicitly laid out my reasons for avoiding evaluating my system with ROGUE scores and I learned really quickly that despite having a terrible scoring metric, the NLP community would rather reject all papers without ROGUE scores reported rather than admit that there is an incredible lack of methods for automatically evaluating summaries or translations.
How many good translation or summarization ideas are not published or utilized just because they don't get high BLEU scores? I bet it's a lot of them...
The author also (perhaps unintentionally) shows a great example of being unable to tell whether a translation is excellent or awful.
猫坐在垫子上 is not a great correspondence to either reference sentence. Going purely by grammar, the equivalent of "the cat is on the mat" would be 猫在垫子上 (note the missing 坐, which is the verb "to sit", not present in the english references), and "there is a cat on the mat" would be 垫子上有猫 [literally "the top of the mat has a cat"].
猫坐在垫子上 could be a good translation of "the cat is on the mat". But it could also be a good translation of "cats sit on mats". Those two sentences are radically different in English; to judge the translation, you need to be aware of whether e.g. someone just asked "where's the cat?" or "what do cats do?"
You just don't see these things reported in the literature because 1. human evals are a pain to run and can be expensive 2. researchers need to publish or perish.
In this post, the 0 example for a good translation arise because the question is not formulated in the standard English way which use sujet/verb inversion. This problem is easily solved by extending the corpus with more diverse way to formulate questions. It’s also notable than the way a question would be translated in Chinese would also similarly be affected by the construct used by the MT system and the constructs present in the corpus. Instead of ditching the scores completely, improving the corpus seems to be a more productive approach, which would also benefit other researchers and the field in the long term.
sacrebleu package helps to standardize BLEU computation to make the comparison of model performance easier. https://github.com/mjpost/sacrebleu