We deal with these hairy problems a lot too. Even before ML proper, getting the definitions right is very tricky. Should a comment that's polite and positive on its own, but supportive of a toxic parent post ("I love Nazis!" -> "I agree!"), be treated as toxic? Is "toxic counterspeech" equally toxic? What about a comedian making fun of an actor's nose, and does it depend on whether the joke is to that actor directly vs. merely referencing them in the third person? etc.
For these reasons, I actually personally like the fact that the Jigsaw annotation guidelines are very high level (as opposed to long and prescriptive) -- it lets the data capture the spectrum of "human preferences" on its own (at least, it does when you can trust that the annotators are able and trying to do a good job).