So wouldn't any algorithm always have some biased in it?
Sorry if this comes off as a stupid question or is unclear.
So wouldn't any algorithm always have some biased in it?
Sorry if this comes off as a stupid question or is unclear.
Any sufficiently complex programming project will end up reflecting some assumptions and biases on the behalf of the programmer. As more programmers contribute to a project, it won't reflect an individual's biases, but will still reflect those programmers as a group.
Can you break this down more? It doesn't make sense to me. If you're writing a machine learning application to take a dataset and match future inputs to past results I don't see how these biases can sneak into the program.
Unless the programmers are changing the datasets then I don't see how this makes sense.
Arrest statistics in the US are heavily skewed by race. If you were to take a dataset of all US arrests between 1900 and 2000 and ask which populations are most likely to "commit crimes" (i.e., be arrested), you'd get racially biased model without recording why it's biased (discrimination, enduring poverty, minimum sentencing, &c).
Why the hell would those be factors? The factors of the case should be things that actually matter in the case.
A rich person who murders someone should get the same sentence as a poor person who murders someone.
A black person who murders someone should get the same sentence as a white person who murders someone.
Adding those factors would be insane in the first place. If you're adding crazy things like that you might as well add factors like "Can Juggle" and "Can Burp the Alphabet" because things like that should have just as much to do with sentencing as what kind of car you drive or where you live.
Model of car is hard to say that it's useful for predictive value, but home address (=which neighborhood did you grow up in, e.g., are you in Cabrini-Green or are you Gold Coast?) and annual income (poor people are more likely to commit crimes than rich people) definitely are going to be queried inputs. Particularly if the question is things like "what should I set bail at?"
Past behavior is a good predictor of future behavior in my experience. That's more than doubled when it's recurring behavior.
> home address (=which neighborhood did you grow up in, e.g., are you in Cabrini-Green or are you Gold Coast?) and annual income (poor people are more likely to commit crimes than rich people) definitely are going to be queried inputs
Why?
> Particularly if the question is things like "what should I set bail at?"
That's still the job of the judge and there is law at hand for how those values are calculated. The computers won't be involved with that (I'd assume) just sentencing based on previous case decisions.
That question includes the past behavior of a lot more people than just the person with a criminal conviction.
> recurring behavior
Institutional racism has recurred a lot in the justice system.
The code that is going to handle the data is also biased in some direction. Its actions is value-based, coming from that of the creators of the algorithm. This is why all code must be open and publicly auditable if used in law enforcement, health care and insurance, etc.
Here's a simple example. Say you're trying to come up with an algorithm that decides whether articles in a data set are "fake news" (topical, I know). We have to tell the algorithm whether a given article in the training set is fake or legitimate, otherwise how would it know? Clearly this will reflect the views of whoever is tagging the articles. When we run the model on a training set we need to score how well it did, again this will reflect the opinion of the person doing the scoring.
For a real example: https://mathbabe.org/2016/05/12/algorithms-are-as-biased-as-....
"*Say you're trying to come up with an algorithm that decides whether articles in a data set are "fake news"
(topical, I know).*"
Current AI cannot do this. This would take finding sources, pulling data out of those sources, cross referencing multiple sources, and recurring for those articles to a certain depth.That's not a good facsimile for deciding sentencing. Sentencing is more like a linear regression classification. You have a history of previous cases where the defendant was found guilty. You then have a pile of factors that played into the judge's decision for sentencing. For example:
* If they meant to do it
* If they feel bad about doing it
* If they did do it (Beyond a reasonable doubt)
* If they have done it before
* What severity this crime is
* ... etc
The judge then uses their experience in law and previous case law as well as statues to find a proper punishment. This is in the form of: * Time served
* Fines
* Privileges revoked
* Community Service
This would then be fed into a classification engine. You leave all of the existing infrastructure in place (Judge, Jury, Lawers) and just use their decision as input into the sentencing.Deciding the validity of claims is not within the scope of modern day machine learning (as of 2017). Classification engines are very much in the scope of machine learning of today.
I don't see how case factors could be biased. I don't see how historical cases (when stripped of all identifying information) could be biased. I don't see why a system like this would be bad.
All treatment of everyone would converge into a uniform handling of cases.
Sure. I'm not saying my example is the smartest (although there are people trying to use ML to do this to be fair - http://www.fakenewschallenge.org/) and I was unaware of your experience with ML, so I was working under the assumption you had no experience with it and trying to pick a simple example (even if it's dumb) to explain how bias can sneak into models in general, rather than in the specific case of sentencing criminals.
Let's loop back to what you originally said:
> If you're writing a machine learning application to take a dataset and match future inputs to past results I don't see how these biases can sneak into the program
You then go on to describe a number of factors that you think should go into sentencing models that leave plenty of scope for bias:
* "If they meant to do it" - this is a judgement made by a person and clearly reflects the view of the person making the decision.
* If they feel bad about doing it - again, someone has to judge whether someone is legitimately remorseful or is trying to pretend they are to get themselves a lighter sentence.
* "If they have done it before" - This will reflect things like policing tactics. For instance poorer areas might be subject to higher rates of policing especially in areas adopting the broken windows theory of policing (https://en.wikipedia.org/wiki/Broken_windows_theory#New_York...) and often in these areas petty crimes are cracked down on more frequently. This means that people are more likely to have run ins with the law, meaning they're less likely to get jobs due to convictions showing up in background checks, which in turn increases their likelihood to reoffend.
* What severity this crime is - I'm not sure what you mean by this. Do you mean e.g. murder being more severe than petty theft, or things like how severe an assault was? I'm assuming the latter since the former is often just covered by things like sentencing guidelines anyway. If someone commits an assault, then how do you rate this in a way that a model can understand? How do you ensure consistency across different cases and judges?
At any rate, the point of these models is usually to remove the biases that judges might have about people of certain backgrounds from sentencing guidelines and produce a score that informs the likelihood of the convict reoffending (I believe), so your proposal isn't how this works in practice. In practice they're trying to avoid exactly these kinds of subjective assessments you proposed and replace them with supposedly objective predictors for the likelihood of the person in question to reoffend. From the article linked from the post: https://www.propublica.org/article/machine-bias-risk-assessm...
> Northpointe’s software is among the most widely used assessment tools in the country. The company does not publicly disclose the calculations used to arrive at defendants’ risk scores, so it is not possible for either defendants or the public to see what might be driving the disparity. (On Sunday, Northpointe gave ProPublica the basics of its future-crime formula — which includes factors such as education levels, and whether a defendant has a job. It did not share the specific calculations, which it said are proprietary.)
> Northpointe’s core product is a set of scores derived from 137 questions that are either answered by defendants or pulled from criminal records. Race is not one of the questions. The survey asks defendants such things as: “Was one of your parents ever sent to jail or prison?” “How many of your friends/acquaintances are taking drugs illegally?” and “How often did you get in fights while at school?” The questionnaire also asks people to agree or disagree with statements such as “A hungry person has a right to steal” and “If people make me angry or lose my temper, I can be dangerous.”
Given that independent research seems to confirm that the company's model seems to favour higher sentences for people of colour, it's pretty clear from that description where biases could sneak in to the model, I hope?
Source: Flores, Bechtel, Lowencamp; Federal Probation Journal, September 2016, "False Positives, False Negatives, and False Analyses: A Rejoinder to “Machine Bias: There’s Software Used Across the Country to Predict Future Criminals. And it’s Biased Against Blacks.”", URL http://www.uscourts.gov/statistics-reports/publications/fede...
In fact the ProPublica analysis was so poorly done that the authors of the above study wrote in the conclusion:
> "It is noteworthy that the ProPublica code of ethics advises investigative journalists that "when in doubt, ask" numerous times. We feel that Larson et al.'s (2016) omissions and mistakes could have been avoided had they just asked. Perhaps they might have even asked...a criminologist? We certainly respect the mission of ProPublica, which is to "practice and promote investigative journalism in the public interest." However, we also feel that the journalists at ProPublica strayed from their own code of ethics in that they did not present the facts accurately, their presentation of the existing literature was incomplete, and they failed to "ask." While we aren’t inferring that they had an agenda in writing their story, we believe that they are better equipped to report the research news, rather than attempt to make the research news."
What I find remarkable is that in the ongoing coverage ProPublica has published on this subject in December 2016 they interviewed a bunch of more people, but none of the folks that have criticized their analysis (published in September). Make of that what you will.
I don't want to end up defending the methods ProPublica have used as I am certainly not qualified to do that and have no skin in this game anyway. I posted here initially in response to a very general question about bias in models, and I'd rather not be drawn into a lengthy discussion about this specific piece.
However, I do have one or two issues with the conclusions you seem to be drawing in your comment:
> What I find remarkable is that in the ongoing coverage ProPublica has published on this subject in December 2016 they interviewed a bunch of more people, but none of the folks that have criticized their analysis (published in September). Make of that what you will
I'm not sure it's possible to conclude anything from that actually. There could be plenty of non-nefarious reasons for the omission. For instance, one individual they cite in the follow up review [1] has written a paper citing the paper you linked to showing that "the differences in false positive and false negative rates cited as evidence of racial bias in the ProPublica article are a direct consequence of applying an instrument that is free from predictive bias to a population in which recidivism prevalence differs across groups". [2] The Flores et al paper seems to claim that showing that predictive bias does not exist is enough, which it would seem is not the case. If racial bias might appear anyway in the situations in which the model is often applied in reality, then perhaps the ProPublica authors felt that the paper cited below [2] adequately addressed the criticism of the paper you cited and decided not to reference the FPJ article for reasons of clarity in their follow up? I think discounting their work because of a single omission would be throwing the baby out with the bathwater.
The ProPublica authors cite plenty of other research in the area in their follow ups. Sure, this is all largely in agreement with their conclusions or go further, but does this matter unless those publications are incorrect? The ProPublica authors are writing for a news publication not an academic journal and are therefore not obligated to cite every relevant publication when they're publishing. So long as they can do this without forcing a conclusion then I don't see the problem. Perhaps they deliberately ignored the paper. Who knows?
[1] https://www.propublica.org/article/bias-in-criminal-risk-sco...
http://www.sciencemag.org/news/2017/04/even-artificial-intel... And the paper http://opus.bath.ac.uk/55288/
Assuming I'm using some actual machine learning model and not my own hand-coded finite state machine, I just don't see how, say, k-means clustering could be biased (unless there was a bug in the programmer's implementation?).
Machine learning generates models based off of input data. Machine learning, principally, will not be [biased or] exhibit any direct bias. [It's math, on its own it's not the problem.]
The generated model is then incorporated into an algorithm for making decisions or categorizing objects. If the original data contains biases, the final model will be biased.
[EDIT: Forgot two words, added a sentence]
In the case of machine learning techniques that employ a learning algorithm (e.g. SGD) and a separate classification algorithm (e.g. forward prop), the situation is no different. Gradient descent in itself doesn't take "unbiased" data and pollute it with "bias" -- neither does, forward propagation.
These algorithms are merely functions that are _parametrized_ by a potentially biased learned model (which really entirely amounts to having biased data).
I recommend taking a look at Stanford's Intro to Machine Learning CS 229 course notes for a brief overview of these concepts. I can't speak anymore highly of the material in this course; it's by far the best out there, and it's how I learned ML.
"The generated model is then incorporated into an algorithm for making decisions or categorizing objects. If the original data contains biases, the final model will be biased."
Right, the "final model" (which comes directly from the data) will be biased, but not the ALGORITHM (the "math," as you call it).
For example, I can't make an unbiased data set "biased" by using linear regression instead of logistic regression. It would be quite a tremendous discovery if you somehow figured out a way SGD in itself was biased, but the burden of proof on such an extraordinary claim would be on you.
So like I said in my parent post, I can see how the data can be biased, but I'm struggling to find any biases in the algorithms themselves (assuming they are properly implemented and not an Underhanded C Contest [1] entry).
Granted, if we're not actually using ML, and we're using something like a hardcoded decision tree or FSM instead (it was pretty strongly implied that we weren't though because we were talking about data-driven algorithms), then sure, that could very easily be a biased algorithm, but I've maintained that all along.
That's correct, but the decision to use gradient descent at all is a huge bias. The mean, median, and mode all describe a dataset's central tendency, but their values can be very different.
Here's a thought experiment for you then: Can you give an example of how one's decision to use gradient descent might impact a data-driven sentencing algorithm?
Or just in general? In contemporary machine learning, it is true that choosing "hyperparameters" is still pretty intuition-guided and based on trial-and-error (even "guessing" if you will).
I wouldn't think those choices could be made maliciously, or even in a "biased" way (okay, _maybe_ I could choose comically out-of-range, nonsensical values -- but that's really a stretch).
If anything, in a hypothetical world, I would feel comfortable letting the accused choose the provably-convergent machine learning algorithm and hyperparameter values of their wishes, and then that would be the sentencing model that is used to determine their fate. As you can see, I claim that their choices wouldn't be able to "bias" anything at all.
As has been true all along, I can certainly see how the _dataset_ could be biased, but I claim that the choice of hyperparameters (and even the algorithm itself) is quite a bit more "immune" to human biases. For one thing, we aren't even capable of fully understanding the learned models produced by state-of-the-art ML algorithms; they are effectively "black boxes."
Despite all of the downvotes, I am genuinely interested in continuing this discussion, and I am curious what you have to say so I will continue to spend more of my karma expressing and arguing for my (unpopular) views.
If these algorithms become widely used and thus influence important outcomes, then people who like to control power will seek to obtain control of the algorithms, either to protect their interests or to grow their own influence.
As possible examples, politicians who see an algorithm apolitically jail a powerful person might pass a law saying that public service should be weighed more heavily. Law enforcement might want exceptions or special weightings for police officer defendants, or for police officer-given testimony. Powerful social activists will buy the algorithm developers (do you know who is writing these algorithms even now? I don't) as a way of influencing society - the Koch brothers, for example, invest in other areas, such as academia and down-ballot elections (e.g., secretary of state in U.S. states) for society-changing purposes. Special code might even generate special outcomes for particular individuals - who will ever see the code? Who will even ask this question after the sentencing?
If you think such corruption couldn't happen, look around. Brazen corruption happens all the time, especially in state and local government. This will be harder to detect, and will be covered by the excuse that non-technical people will believe: I didn't do it; it was the algorithm! It must be objective!
I think the algorithms ultimately will be a step backward, reducing transparency by hiding the bias in a software development process, in the back room of a software firm, and in a mountain of code. At least legislation is published and voted on in the open.