One, if datasets are biased --- if you build your system to only work on white males --- then it may have suboptimal results for other groups. This is a common problem: you use your company's faces, or college students enrolling in data gathering exercises, etc, who are not representative of the population at large. We can fix this by being careful about dataset bias.
But the second issue hits right at the heart of a major societal problem/debate. When we use AI to make decisions about people, will the system become racist--- even with representative datasets? If you train something to predict, say, odds to default on a loan-- will it figure out things that correlate to race and be making mostly racial decisions? Different races do have different default rates, but we've decided as a society that it is unfair to use race to determine an individual's probability of default. But if we choose things that are correlates of both default rates and race, when is that fair and when is it just veiled racism (redlining)? What things are measuring a causal relationship and what things are just racism in disguise?
This second problem is much worse with ML, because we have the ability to accept a whole bunch more things into our models and explaining the rationale of why decision are made is much harder.
And of course, the first problem-- both bias in datasets during use and biases during research and development -- makes the second problem worse. It can't even necessarily be addressed by broadening the dataset and retraining: if, in this case, you do your research and training with just white faces, and report positive results, it may not generalize to work as well for everyone with a broader dataset.