For starters, you can't unless you actually include the context in the training set.
Then again, a lot of humans won't successfully manage that either...
For starters, you can't unless you actually include the context in the training set.
Then again, a lot of humans won't successfully manage that either...
It's worse than that. The context doesn't just have to be in the training set. It also has to be available on the other end, when you're using the model to make a determination. And it isn't.
Which is why even human censors can't get it right. The parties communicating can, and almost always do, have external shared state that The Decider doesn't. Which gives the same words different meaning.
Imagine hearing an inside joke you're on the outside of and then being asked to adjudicate whether it was offensive.
E.g. if judging tweets for example you'd presumably do a lot better if you evaluated the tweets in context of followers and in context of past tweets - both recent and the accounts history. E.g. an account that posts white supremacy tweets regularly is likely to mean something entirely different if RT'ing a BLM tweet with "that is fucking awesome" than what someone with #BLM in their profile is likely to mean.
Part of the problem is that they're throwing away a huge amount of the state they do in fact have.
But I absolutely agree it's in general an unsolvable issue - I pointed to a story from my childhood elsewhere showing how trivially people cause problems for "The Decider": A childhood friend being forced not to call his brother names and switching to using names of cheeses.
Is "Edam" an insult or a food preference? You can't know without context.
And it needs to be very local context, because humans very quickly pick up on when you just start using a word to mean something else, so it doesn't even need to be any shared external state about the word, just a shared understanding of where the receiver might expect an insult coupled with an unexpected response that will then easily get labelled an insult. That unexpected term might well in itself be positive if the receiver expects criticism. "Awesome" and "I love it!" are perfectly good insults when the other party has just told you something where the appropriate response would be negative, for example.
Must never use followers. Or anything else the account holder has no control over. Otherwise you'll end up with the same kind of SEO problems where competitors and foreign governments create fake accounts to follow or interact with disfavored ones and destroy their algorithmic reputation.
Guilt by association in general is malice. Someone who regularly interacts with white supremacists might be a white supremacist -- and maybe 90% of them are -- but it could also be someone criticizing, mocking or debunking them. And if it gets out that the algorithm is penalizing people for doing that, they'll stop. Which is very bad.
I specifically follow a few locals that I would often notice upset with the same politicians that I was, but for diametrically different reasons. Even though I disagree with them on most everything, I find value in having a few of those voices on my feed visibly attached to a consistent person. It helps me to resist seeing the "other side" as just an impersonal sea of voices.
100%. It's crazy how often context is missing from datasets! We wrote a separate blog post recently about this too: https://www.surgehq.ai/blog/why-context-aware-datasets-are-c...