I'd rather have people too concerned about ethics than not enough.
Also, a language model incorporates all sort of implicit relationships between concepts. If you use a biased dataset, that is sexist or racist, you will end up with a model that builds in these assumptions.