I had the same experience when trying to do NER on customer support requests. My model performed great for research datasets but it was mediocre at best for my own dataset. Do you have any suggestions on how to achieve better results in domains where mistakes, bad punctuation, etc are common?