A lot of recent research uses RL to "fine-tune" NLP models. A practical example would be Google's recently announced Machine Translation System (https://arxiv.org/abs/1609.08144). It uses RL to directly optimize BLEU scores on translated sentences.
You'll find similar applications in state-of-the art models for chatbots for example. Though I agree, "widely used" may be somewhat of an overstatement. But it's becoming more common.
On a side note, I actually think RL makes a lot of sense for many NLP problems and it would be super interesting to build a pure RL approach to language modeling or translation. Nobody has managed to do that quite yet.