For example, try using named entity models trained on CoNLL (newspaper articles) on free text (e.g. tweets, or text from application forms), you generally get pretty bad results. When the domain is different, I've even seen them screw up basic things like times and dates, where regexes will suffice. If you're using it for newspaper articles you're sorted, if you're not, the performance metrics here are probably not all that meaningful.