Well, BLEU is non-differentiable and not decomposable over sequence of translation decisions. Yet I wouldn't call methods reinforcement learning because loss is tricky.
But yeah, I guess there's more to it than meets the eye.
But yeah, I guess there's more to it than meets the eye.