I'm leaning against the user flagging model of content moderation because it keeps the incentive for bad actors to try to game the system. Bad content, even briefly included can still do a lot of damage. It's like spam email -- even though something like 99.9% gets filtered out, that 0.1% is still valuable enough for spam senders to keep trying. If the problem of spam is just a numbers game, then bad actors will just increase the volume to compensate.