Even if you think the example presented here isn't a big deal, it still exposes a larger issue in these large language models. These models learn from historical data and that history may not be the behavior we want in the future.
Is your point that there be disagreement on ML policy? I don't think that makes any difference to the fundamental research problem.