Yeah, I think this is part of the problem. Is large-scale, low-quality data good? Sometimes it is (depending on the tradeoff), but from a model performance perspective, it's often more effective to get smaller amounts of higher-quality data instead.
Hopefully people also don't need to be at the level of a professional linguist to label messages like "this is fucking awesome" correctly!
And great point on context. For example, the GoEmotions dataset didn't present labelers with the actual post or subreddit the message came from -- just the text itself. That makes it really difficult to label something like "his traps hide the fucking sun"! But once you see the comment in its original context https://www.reddit.com/r/nattyorjuice/comments/aee3wx/olympi..., and know that it's in the /r/nattyorjuice bodybuilding subreddit, it's much easier to realize that this is talking about someone's large muscles.