One paper that really hammers this point: Inherent Trade-Offs in Algorithmic Fairness [1]
The example they focus on is different but the general principles and takeaways are very powerful and applicable to all classification problems including content flagging.