Sorry, but I think it's a really dark road to have a tool determine and reify what's "harmful content" and then characterize a question through that lens in its response. It's a kind of cultural hegemony and we need to be really careful about embedding that into these $MM systems if there are only going to be a handful of them.
It's very easy to point to straightforward, contemporary examples like illicit self-harm or bomb-making and say that these are plainly harmful and through those justify the system behavior -- but that's blind to the innumerable topics that live on the edge of cultural difference (by time, geography, ethnicity, etc).
Can you imagine if these were a product of 1980's AI research and codified some of that time's widespread ideas about sexual orientation or even atheism? "I’m really sorry to hear that you’re feeling this way, but I can’t provide the help that you need. It’s important to talk to someone who can, though, such as a mental health professional or a trusted person in your life."
What we should probably be doing is recognizing that universal general assistance is a poor fit for these tools since there isn't a universal general culture that they can align with. Instead, we should look towards fine-tuning to make them purpose based ("Sir, this is a Wendy's") or make them sufficiently open and re-deployable so that cultural norms can be fine-tuned over a nonjudgmental baseline.
Insofar as "AI alignment" pretends that we all have the same ethical orientation and that the AI should be made to align with it, it's reinvigorating some very dark ideas from days of empire and colonialism. The fact is that humans aren't ethically aligned with each other, and aligning centralized AI with some particular community is way of projecting that community's values on everybody else.