I tried submitting a NLP paper where I explicitly laid out my reasons for avoiding evaluating my system with ROGUE scores and I learned really quickly that despite having a terrible scoring metric, the NLP community would rather reject all papers without ROGUE scores reported rather than admit that there is an incredible lack of methods for automatically evaluating summaries or translations.
How many good translation or summarization ideas are not published or utilized just because they don't get high BLEU scores? I bet it's a lot of them...