The fact that he does this for free is also concerning, primarily as I doubt this has any level of auditing behind it. The only thing I agree with him on is that black box models are even worse as they have even worse audit issues. Given the complexities in making these predictions and the potentially life long impact they might have, there is such a desperately strong need for these systems to have audit guarantees. It's noted that he supposedly shares the code for his systems - if so, I'd love to see it? Is it just shared with the relevant governmental departments who likely have no ability to audit such models? Has it been audited?
Would you trust mission critical code that didn't have some level of unit testing? Some level of code review? No? Then why would you potentially destructively change someone's life based on that same level of quality?
> "[How risk scores are impacted by race] has not been analyzed yet," she said. "However, it needs to be noted that parole is very different than sentencing. The board is not determining guilt or innocence. We are looking at risk."
What? Seriously? Not analyzed? The other worrying assumption is that it isn't used in sentencing. People have a tendency to seek out and misuse information even if they're told not to. This was specifically noted in another article on the misuse of Compas, the black box system. Deciding on parole also doesn't mean you can avoid analyzing bias. If you're denying parole for specific people algorithmically, that can still be insanely destructive.
> Berk readily acknowledges this as a concern, then quickly dismisses it. Race isn’t an input in any of his systems, and he says his own research has shown his algorithms produce similar risk scores regardless of race.
There are so many proxies for race within the feature set. It's touched on lightly in the article - location, number of arrests, etc - but it gets even more complex when you allow a sufficiently complex machine learning model access to "innocuous" features. Specific ML systems ("deep") can infer hidden variables such as race. Even location is a brilliant proxy for race as seen in redlining[1]. It does appear from his publications that they're shallow models - namely random forests, logistic regression, and boosting[2][3][4].
FOR THE LOVE OF EVERYTHING THAT'S HOLY STOP THROWING MACHINE LEARNING AT EVERYTHING. Think it through. Please. Please please please. I am a big believer that machine learning can enable wonderful things - but it could also enable a destructive feedback loop in so many systems.
Resume screening, credit card applications, parole risk classification, ... This is just the tip of the iceberg of potential misuses for machine learning.
Edit: I am literally physically feeling ill. He uses logistic regression, random forests, boosting ... standard machine learning algorithms. Fine. Okay ... but you now think the algorithms that might get you okay results on Kaggle competitions can be used to predict a child's future crimes?!?! WTF. What. The actual. ^^^^.
Anyone who even knows the hello world of machine learning would laugh at this if the person saying it wasn't literally supplying information to governmental agencies right now.
I wrote an article last week on "It's ML, not magic"[5] but I didn't think I'd need to cover this level of stupidity.
[1]: https://en.wikipedia.org/wiki/Redlining
[2]: https://books.google.com/books/about/Criminal_Justice_Foreca...
[3]: https://www.semanticscholar.org/paper/Developing-a-Practical...
[4]: https://www.semanticscholar.org/paper/Algorithmic-criminolog...