We've had several "customer feedback / intent / support case analysis" projects in the past. Some for large customers with millions of individual records (Autodesk), where there's the additional challenge of "What should the categories be in the first place? What's in the data?" (discovery).
What we learned is a model trained on one type of feedback will not necessarily perform well on others, because the relevant signals manifest differently across modalities: feedback length / writing style / typos, lexical richness / repetition / boilerplate, OCR noise / how long is the long tail… Your model may learn to pick up on cues that are orthogonal to the sentiment or categorization problem.
This is especially true for black box models (deep learning) where introspection is limited: Did the model learn to rely on syntax? Specific words or character ngrams? Exclamation marks? Something else? Does an Indian-looking name imply sentiment negativity?
Slapping a generic ML technique (Stanford NLP, Naive Bayes, bi-LSTM, whatever) onto a bunch of tokens is a reasonable first step, that's the low-hanging fruit. The tricky part is defining the problem space and the QA process correctly, and managing the devil that comes with the details.