Algorithm can find a neighborhood’s political leanings by its cars
news.stanford.edu
news.stanford.edu
even the existence of this correlation is not obvious
E.g. If you work in construction, or you hunt, you'd be a lot more likely to choose a pickup truck for purely practical reasons.
I have to wonder sometimes if owning a certain kind of car "just because you like it" is as reasonable a reason for ownership as I consider it to be.
How do they ascertain that the cars belong to the suburb? Would a suburb be ranked differently if for a month there were civil works that meant a lot of contractor vehicles etc were parked up? What about if the local Ferrari club decides to go for a cruise and all meet at a coffee shop in the neighbouring suburb every week? What if the suburb was located in a cold region, where it's more practical for most people to drive a SUV in winter?
What about the republicans who drive Civics, and the republicans who drive Audis? What about the democrats that drive Toyota Tacomas, or the democrats who drive Volts?
I mean, what would it make of me? I have a toyota hilux for work, a mx5 (miata) for track days, and a subaru wrx sti for driving to the shops. I wonder what my political leaning would be considered to be?
All machine learning involves a cost function. How close do the results have to be to make a useful prediction in the context.
As soon as you have a researcher make this sort of claim, a politician will inevitably see it as "Don't spend money on census research, do this instead" ...
I presume they looked at more than one neighborhood? Or are you suggesting Ferrari owners have a conspiracy to do this in general?
But to be serious, my point was more about the way that they are collecting their data. If you were focusing on buildings and structures in a suburb to rate demographics, I would think that you would have more accurate results. Cars don't stay still though, and just because it's in a location at a particular time doesn't necessarily mean it lives there. The ferrari club going for coffee is just an example of affluent (types of) cars possibly throwing out results.
Maybe their conspiring is the reason they got their Ferraris in the first place ;-)
Of course you can't assert things on a single individual based on his car. Just give a probability based on all the evidence you have.
But on a great number of people (a neighborhood) it doesn't surprise me that you get pretty good results.
I'm sorry, what? A Bentley is "understated"?
After all, when oil prices rise you'd expect people to choose more efficient cars - not necessarily to change their voting patterns. I can't see any mention in the original paper of how they calibrated for that.
http://www.pnas.org/content/early/2017/11/27/1700035114.full...
Could you point us to whatever sentence in the article that made or even suggested this? I didn't see any part of the article stating that everyone who has a pickup truck is conservative.
"Everyone with a pickup truck is conservative/Republican."
from the statement
"If P(sedans) > P(pickups) then P(precinct voting Democrat) >= 88%."
?
You just learned what it means.
Got a training set which dates to when banks redlined black people? Congratulations, the model just learned not to make loans to black people! And now it's "objective", because no humans are involved in making its decisions!
Assuming, ridiculously conservatively that redlining ended in 1979 any dataset will have 37 non redlined years in it, 1980-2017. Given that the further away in time a datapoint is the less weight it’s given in predicting anything redlining will have bigger all impact on any model.
Asian people in the US on average have higher credit scores than white people. This is not anti-white discrimination because white and Asian are not used to determine credit score, it’s based on on-time payment, capacity used, length of credit history, types of credit used and past credit applications.
There are many wonderful MOOCs and textbooks on ML, AI and statistics. If you want to understand how statistical modelling works I can recommend the Georgia TechX MicroMasters in Analytics.
Don’t listen to statistically illiterate journalists.
Redlining is only one of the many racists housing practices that have existed and continue to exist today.l
> in 1979
Have you even done any research on this topic? From a famous article[1] about the history of racism in housing:
>> In 2011, Bank of America agreed to pay $355 million to settle charges of discrimination against its Countrywide unit. The following year, Wells Fargo settled its discrimination suit for more than $175 million. But the damage had been done. In 2009, half the properties in Baltimore whose owners had been granted loans by Wells Fargo between 2005 and 2008 were vacant; 71 percent of these properties were in predominantly black neighborhoods.
> 37 non redlined years
Those are also years of sub-standard schools, poor job prospects, and (in some cities) toxic levels of lead paint. Growing up in that kind of environment can limit your options severely.
> it is not possible for your model to discriminate against Black people because it doesn’t know if anyone is black.
I always like to assume ignorance over malice, so I will simply suggest that you should do a lot more research before making this kind of claim.
[1] https://www.theatlantic.com/magazine/archive/2014/06/the-cas...
Hadn't heard of it. Sounds catchy though. Could you explain the similarity? (Along with the definition of "bias" you're using?)
Intent is not required.
> only going to have evidence-based biases
That includes historic institutional biases. For example, many US cities are still highly segregated by race. This was done intentionally as a workaround for forced-integration rulings with indirect methods such as redlining, blockbusting, and restrictive covenants. While most of these methods were eventually banned, the results of those methods still exist today[1].
Since school funding usually depends on property taxes, these policies also impacted the school system. It may be a fact that one group of people performs worse on some metric, but how much of that was caused by racist housing practices decades earlier? Unless you're accounting for all of the embedded biases already present in the facts, you are laundering those biases into your results.
[1] https://www.huffingtonpost.com/entry/the-9-most-segregated-c...
Whitewashing the training data to make predictions based on a hypothetical injustice-free world where redlining never happened will just give you garbage predictions.
> and start being facts with unjust causes.
Of course. That means you know the data contains an "unjust" bias. Do you want to account for that bias, or do you want to proceed anyway risking problems (e.g. Simpson's Paradox).
There are more confounding factors in heaven and earth, closeparen, than are dreamt of in your model.
https://mathbabe.org/2017/01/04/recidivism-risk-algorithms-a...
https://mathbabe.org/2013/01/29/bill-gates-is-naive-data-is-...
A risk-management approach to the criminal justice system doesn’t call for fudge factors. It’s fundamentally unacceptable. Risk management is all about using people’s demographics against them (see: all insurance pricing). Wokeness on the part of the risk managers doesn’t fix that. The criminal justice system needs to treat people as individuals and not classes - this requires that it steer clear of evidence-based risk assessment.
Thus type of.machine learning is externalized, automated prejudice.
See also the gay detector.
"We stopped testing the medicine because even though it cured 99% of cases those that weren't cured were offended by the possibility that someone might assume that it would cure them"
Statistical reasoning becomes prejudice when you act on it in a way that is harmful to society or you overgeneralize it. If you try to ban it then you're just avoiding emotionally uncomfortable truths.
Succinct
It's similar to how the distribution of height in men and women is different, ie, women tend to be smaller (IIRC), it has no bearing on the expression of the measures values on an individual.
There is a probability that you are conservative if you have a pickup truck but this doesn't mean you are conservative if you own a pickup truck.