It’s out of the scope of this library to publish all datasets in existence or highlight particular datasets that are relevant to particular societal problems. It’s literally just a few datasets so that you can play around with the ML library without downloading any external datasets. I think it’s fair to allow them to exercise reasonable discretion in their choice of which toy datasets to ship with their ML library.
I.e., you cannot distinguish a 73% black neighborhood from a 53% black neighborhood with this variable.
It's a bizarre variable and I guess I could see purging the column or at least suggesting it not be used, but I don't really understand why you'd delete the rest of the (sample) dataset on this basis.
> I.e., you cannot distinguish a 73% black neighborhood from a 53% black neighborhood with this variable.
Isn't the point of that operation that a 1% black neighborhood and a 99% black neighborhood are both less integrated than a 50% black neighborhood? If you didn't do something like squaring, then wouldn't at least one of the former incorrectly register as more integrated than the latter?
> It's a bizarre variable and I guess I could see purging the column or at least suggesting it not be used, but I don't really understand why you'd delete the rest of the (sample) dataset on this basis.
Agreed.
"Thus, any models trained using this data that do not take special care to process B will learn to use mathematically encoded racism as a factor in house price prediction."
ML models could also learn that the correlation between B and price is negative (i.e., that integration improves house prices). But the critics of the dataset all suggest that B and price are positively correlated.
I think regardless removing this model seems right since it's A a toy, B dated, and C liable to not be seen with the care it needs to be, but say if I'm trying to calculate who will be the next famous actor, is it wrong to include variables to allow the model to pick up on how shallow and say fat phobic society is (or just plain racist again)
The Boston housing prices dataset has an ethical problem: as
investigated in [1], the authors of this dataset engineered a
non-invertible variable "B" assuming that racial self-segregation had a
positive impact on house prices [2]. Furthermore the goal of the
research that led to the creation of this dataset was to study the
impact of air quality but it did not give adequate demonstration of the
validity of this assumption.
The scikit-learn maintainers therefore strongly discourage the use of
this dataset unless the purpose of the code is to study and educate
about ethical issues in data science and machine learning.And the transformation itself is non-invertible, so it's not possible to recover the original values for about 7-8% of the rows in the dataset. The commit diff links to a thorough investigation of the data[1] in which the author takes a crack at linking up the ambiguous rows with the original 1970 Census data that supposedly went in to generating this dataset, and long story short, it looks like the original dataset's authors may have made some errors in their calculations on top of everything else.
1: https://medium.com/@docintangible/racist-data-destruction-11...
It's a ridiculous and offensive premise from my perspective.
It isn't. Bk is "number of blacks in my neighbourhood" as you put it, and the whole point of using B instead of it was so that an all-black neighborhood wouldn't count as more integrated than one with a mix of races.
I wonder if the dataset design makes more sense in the context of Boston in particular: [0].
[0] https://en.wikipedia.org/wiki/Boston_desegregation_busing_cr...
Can you elaborate on the problem you have with this?
(I'm just trying to not guess at your meaning.)
As far as I can tell, it's something like, "The author made a bad variable. Also, the goal was to check air quality but the variable was bad." What does the subject being air quality impact have to do with anything there?
1000(Bk - 0.63)^2 where Bk is the proportion of blacks by town
And not sure how anyone can argue the dataset is worthy of being included. It is pretty offensive and misguided at minimum to argue that having more black people in your neighbourhood will depress housing prices. And for it to be solely because they are black and not to do with a range of other factors e.g. socio-economic.
It is pretty offensive and misguided at minimum to argue that having more black people in your neighbourhood will depress housing prices.
I think that's the wrong lens to look at this through. I'm happy to concede your statement about it being offensive is true (although I think from a purely statistical perspective, correlations with poverty, etc probably make the assumption correct. Before 2015 or so when we all lost it, it would only be racist to say there was a causal relationship between race and price, not a correlation). Anyway, that's all an aside.It's the purging of a dataset, a toy dataset in this context, for a reason of political correctness, that I don't support. If you look hard enough at anything, you can probably find a way to call it racist or some similar slur. If we start applying this lens to tools like scikit learn, we go down a path I don't agree with, that's completely performative in terms of actually addressing any wrongs, and is a continuing distraction from what could be useful work. Debating if and how racist this is is immaterial imo to whether or not we should erase everything doesn't align with modern hypersensitivity about political correctness
So even if you want to ignore the fact that the data was outright used to discriminate in the past, the data itself is actually flawed in several ways...
Why are people singling out this one dataset?
Which is such an insane premise that it's hard to take the rest of your point seriously.
The data is the data. The data isn't suggesting that "having more black people in your neighbourhood will depress housing prices." That's your take on what a racist causal interpretation would look like.
The correlation is very real and turning a blind eye to it is worse: https://www.brookings.edu/testimonies/how-racial-disparities...
> At low to moderate levels of B, an increase in B should have a negative influence on housing value if Blacks are regarded as undesirable neighbors by Whites.
Isn't that true? It's almost a tautology.
(Bad things can be true.)
For example ignores the fact that on the whole black people are significantly poorer.
In Boston the average white family has a net worth of $200k+ while black families in the same city have a net worth of <$10 (that's not a typo, it's less than ten dollars). Poorer people by nature can not afford houses in more expensive neighborhoods, so naturally you have a concentration of black people in poorer neighborhoods... it's not because white people find black neighbors undesirable, it's that black people are disproportionately poorer.
This kind of oversimplification tends to perpetuate a lot of negative stereotypes about black people while hand-waiving away the chronic issues black people are faced with that creates this kind of disparity.
> At low to moderate levels of B, an increase in B should have a negative influence on housing value if Blacks are regarded as undesirable neighbors by Whites.
Either way, it's a dramatic oversimplification. Black people are also significantly poorer than white people, so it's also likely that black people can only afford to move into a neighborhood as property values decrease... yet the people who collated this data chose to make a different assumption.
The more it correlates with a lower price, the more race issues.
It's seems like a weird choice. If there are 62% or 64% then the feature will yield the same value.
I think it would make more sense just to include percentage of households where at least one member of the household has a certain ethnicity.
I don't think it's a offensive to analyze demographic information in the aggregate. In fact this happens all the time.
b) It is not scikit-learn's responsibility to alter third party datasets.
IIUC, you're arguing that "whole-town" level aggregation is misleading. So if we get more granular, we could do it my neighborhood, street, building, apartment/car/shelter, bedroom, bed, bed @ time of day, etc.
Any one of those aggregation levels could hide interesting distinctions that could be made if only the data were reported with even more granularity.
So are you arguing against aggregation in general? Or just whole-town aggregation specifically?