[1] https://noisy-text.github.io/2017/emerging-rare-entities.htm...
Do more clustering.
Label more training data.
Strip out more garbage.
Label more training data.
PS you can get an idea of how much value additional training data will give you by training models on various subsets of your dataset (e.g. 10%, 20%...), evaluating them against the same test dataset, and plotting the results.
If one is trained on screwing light bulbs that training would not be very helpful in composing music. If there is some common structure between the train and the test scenarios there would be some point in learning that from the default train set. Then you use your domain specific training set to unlearn the things that do not apply and learn the other things that do. As long as there is something worth learning from the default train set it will be of some use.