In the days when Sussman was a novice, Minsky once
came to him as he sat hacking at the PDP-6.
"What are you doing?", asked Minsky.
"I am training a randomly wired neural net to play Tic-Tac-Toe" Sussman replied.
"Why is the net wired randomly?", asked Minsky.
"I do not want it to have any preconceptions of how to play", Sussman said.
Minsky then shut his eyes.
"Why do you close your eyes?", Sussman asked his teacher.
"So that the room will be empty."
At that moment, Sussman was enlightened.
(from: http://catb.org/jargon/html/koans.html#id3141241 )There is always bias (and other types of error) in data. There is even bias in the choice of which data to use and the type of analysis to perform.
If you think that this isn't a problem, you should really read about practices like "redlining"[1]. For many decades segregation was (and still is) enforced by opaque "loan approval" methods that just happened to always deny loans to blacks.
[1] http://www.theatlantic.com/magazine/archive/2014/06/the-case...
Thus, unless the goal is to perpetuate the effects of discrimination, algorithms need to be carefully set up and trained to understand that "this isn't how things are supposed to be".
Now, take that model and apply a really crude analysis. You would conclude that there is five times more crime in the more heavily patrolled neighborhood. Your algorithm, applied to policy, suggests that crime could be controlled better by patrolling the "bad" neighborhood ten times as much, rather than five times as much.
Now, consider the social result. The "high crime neighborhood" then sees its property values drop, and more homes sold into the rental market. This attracts poorer, rougher people who can't afford the "good" neighborhood. Now crime actually goes up, reinforcing the idea that the bad neighborhood is high-crime.
Now, imagine we've been doing this to those two neighborhoods for generations. What are the likely results of your algorithm? There's your uncomfortable truth.