It indicates that the model is learning a good likelihood function, and that it has still not trained to convergence which makes the samples have lower likelihood scores than real text. That's all that's needed to make it a good recognizer.
No comments yet.