Boston housing price dataset was removed from scikit-learn 1.2
github.com
github.com
Contrary to the cited Medium post, including race, integration, or racism-related factors as predictors of housing prices doesn't imply that it's okay for those drivers to exist in reality, or that the model is suitable for some real-world deployment where it might affect real prices. That is not the only use case; such a model could just as easily be used to imply that those factors are bad. Such a model could play a role in a disparate impact case fighting against racism.
I don't love how the dataset creators only included B and not the underlying untransformed value, and I agree that it's based on a questionable theory about how integration affects housing prices. These issues could be taken as sufficient cause to stop using the dataset, especially when better alternatives are available. But calling them ethical issues seems either puzzling or wrong. Problematizing something should ideally come with an accessible public explanation of why it is problematic. Ethics should not be obscure.
There's nothing baffling about it.
For what it's worth, another commenter posted this and I found it much clearer, if still imperfect: https://fairlearn.org/main/user_guide/datasets/boston_housin...
The FairLearn authors here explain, like I did above, that the ethical issues depend on use case, and are not inherent to the data or the dataset developers' choice of transformation for the B variable.
If the dataset gave the raw features it would be better at least
Edit: D'oh 62% and 64% black neighborhoods get the same B. Couple glasses of wine in already, math is hard...
It’s out of the scope of this library to publish all datasets in existence or highlight particular datasets that are relevant to particular societal problems. It’s literally just a few datasets so that you can play around with the ML library without downloading any external datasets. I think it’s fair to allow them to exercise reasonable discretion in their choice of which toy datasets to ship with their ML library.
I.e., you cannot distinguish a 73% black neighborhood from a 53% black neighborhood with this variable.
It's a bizarre variable and I guess I could see purging the column or at least suggesting it not be used, but I don't really understand why you'd delete the rest of the (sample) dataset on this basis.
> I.e., you cannot distinguish a 73% black neighborhood from a 53% black neighborhood with this variable.
Isn't the point of that operation that a 1% black neighborhood and a 99% black neighborhood are both less integrated than a 50% black neighborhood? If you didn't do something like squaring, then wouldn't at least one of the former incorrectly register as more integrated than the latter?
> It's a bizarre variable and I guess I could see purging the column or at least suggesting it not be used, but I don't really understand why you'd delete the rest of the (sample) dataset on this basis.
Agreed.
"Thus, any models trained using this data that do not take special care to process B will learn to use mathematically encoded racism as a factor in house price prediction."
ML models could also learn that the correlation between B and price is negative (i.e., that integration improves house prices). But the critics of the dataset all suggest that B and price are positively correlated.
I think regardless removing this model seems right since it's A a toy, B dated, and C liable to not be seen with the care it needs to be, but say if I'm trying to calculate who will be the next famous actor, is it wrong to include variables to allow the model to pick up on how shallow and say fat phobic society is (or just plain racist again)
The Boston housing prices dataset has an ethical problem: as
investigated in [1], the authors of this dataset engineered a
non-invertible variable "B" assuming that racial self-segregation had a
positive impact on house prices [2]. Furthermore the goal of the
research that led to the creation of this dataset was to study the
impact of air quality but it did not give adequate demonstration of the
validity of this assumption.
The scikit-learn maintainers therefore strongly discourage the use of
this dataset unless the purpose of the code is to study and educate
about ethical issues in data science and machine learning.And the transformation itself is non-invertible, so it's not possible to recover the original values for about 7-8% of the rows in the dataset. The commit diff links to a thorough investigation of the data[1] in which the author takes a crack at linking up the ambiguous rows with the original 1970 Census data that supposedly went in to generating this dataset, and long story short, it looks like the original dataset's authors may have made some errors in their calculations on top of everything else.
1: https://medium.com/@docintangible/racist-data-destruction-11...
It's a ridiculous and offensive premise from my perspective.
It isn't. Bk is "number of blacks in my neighbourhood" as you put it, and the whole point of using B instead of it was so that an all-black neighborhood wouldn't count as more integrated than one with a mix of races.
I wonder if the dataset design makes more sense in the context of Boston in particular: [0].
[0] https://en.wikipedia.org/wiki/Boston_desegregation_busing_cr...
Can you elaborate on the problem you have with this?
(I'm just trying to not guess at your meaning.)
As far as I can tell, it's something like, "The author made a bad variable. Also, the goal was to check air quality but the variable was bad." What does the subject being air quality impact have to do with anything there?
1000(Bk - 0.63)^2 where Bk is the proportion of blacks by town
And not sure how anyone can argue the dataset is worthy of being included. It is pretty offensive and misguided at minimum to argue that having more black people in your neighbourhood will depress housing prices. And for it to be solely because they are black and not to do with a range of other factors e.g. socio-economic.
It is pretty offensive and misguided at minimum to argue that having more black people in your neighbourhood will depress housing prices.
I think that's the wrong lens to look at this through. I'm happy to concede your statement about it being offensive is true (although I think from a purely statistical perspective, correlations with poverty, etc probably make the assumption correct. Before 2015 or so when we all lost it, it would only be racist to say there was a causal relationship between race and price, not a correlation). Anyway, that's all an aside.It's the purging of a dataset, a toy dataset in this context, for a reason of political correctness, that I don't support. If you look hard enough at anything, you can probably find a way to call it racist or some similar slur. If we start applying this lens to tools like scikit learn, we go down a path I don't agree with, that's completely performative in terms of actually addressing any wrongs, and is a continuing distraction from what could be useful work. Debating if and how racist this is is immaterial imo to whether or not we should erase everything doesn't align with modern hypersensitivity about political correctness
So even if you want to ignore the fact that the data was outright used to discriminate in the past, the data itself is actually flawed in several ways...
Why are people singling out this one dataset?
Which is such an insane premise that it's hard to take the rest of your point seriously.
The data is the data. The data isn't suggesting that "having more black people in your neighbourhood will depress housing prices." That's your take on what a racist causal interpretation would look like.
The correlation is very real and turning a blind eye to it is worse: https://www.brookings.edu/testimonies/how-racial-disparities...
> At low to moderate levels of B, an increase in B should have a negative influence on housing value if Blacks are regarded as undesirable neighbors by Whites.
Isn't that true? It's almost a tautology.
(Bad things can be true.)
For example ignores the fact that on the whole black people are significantly poorer.
In Boston the average white family has a net worth of $200k+ while black families in the same city have a net worth of <$10 (that's not a typo, it's less than ten dollars). Poorer people by nature can not afford houses in more expensive neighborhoods, so naturally you have a concentration of black people in poorer neighborhoods... it's not because white people find black neighbors undesirable, it's that black people are disproportionately poorer.
This kind of oversimplification tends to perpetuate a lot of negative stereotypes about black people while hand-waiving away the chronic issues black people are faced with that creates this kind of disparity.
> At low to moderate levels of B, an increase in B should have a negative influence on housing value if Blacks are regarded as undesirable neighbors by Whites.
Either way, it's a dramatic oversimplification. Black people are also significantly poorer than white people, so it's also likely that black people can only afford to move into a neighborhood as property values decrease... yet the people who collated this data chose to make a different assumption.
The more it correlates with a lower price, the more race issues.
It's seems like a weird choice. If there are 62% or 64% then the feature will yield the same value.
I think it would make more sense just to include percentage of households where at least one member of the household has a certain ethnicity.
I don't think it's a offensive to analyze demographic information in the aggregate. In fact this happens all the time.
b) It is not scikit-learn's responsibility to alter third party datasets.
IIUC, you're arguing that "whole-town" level aggregation is misleading. So if we get more granular, we could do it my neighborhood, street, building, apartment/car/shelter, bedroom, bed, bed @ time of day, etc.
Any one of those aggregation levels could hide interesting distinctions that could be made if only the data were reported with even more granularity.
So are you arguing against aggregation in general? Or just whole-town aggregation specifically?
Think of all the children books that get rewritten. Read the new ones to your children and discuss the old ones when they are teenagers. I would have preferred if sklearn contributors had done the same and simply revised the description as opposed to removing the dataset.
EDIT: changed "banning" to "removing" the dataset
That I think gets us a full ELI5, though I agree the common-sense cutoff is subjective.
(And I say this in a spirit of co-collaboration -- I'm fascinated by the problem of ELI5 something like this, as I wind up having these conversations ad-hoc in-the-wild with family and friends, as I work in the field. Finding simple language just is progress)
Have you ever talked to a five-year-old? They're very interested but ignorant and have minute attention spans. Have some self-respect, put on your big boy pants and at least make an effort to understand the adult version and ask some specifics about what you don't understand.
What you're asking for is a full on Malcom Gladwell style essay. Not only that, but all the points are discussed in the linked blog post. If one can't even be asked to put in the effort to read it, isn't it a bit presumptuous to expect anyone to summarize for you in elaborate detail!?
ELI5 would be: The authors wanted to find out if housing prices have something to do with bad air. They thought some people unfairly make homes cheaper if black people live in them. And that in some cases more expensive. Instead of including data about how many black people live in a neighborhood he included a value that supports his opinion and can't be checked. Because scientists aren't supposed to base their conclusions on gut feelings, this makes the sample data an example for problems that arise when working with data.
Five year olds are not particularly intellectual and wouldn't enjoy the discussion of biased models with feature-invertible data!
'Although models that make use of non-invertible, artifical features may still be biased, it's less likely. Using only invertible features eliminates some risks of bias, but not others.'
At this point I think one might have to explain the Naturalistic Fallacy (which accounts for much of the remaining possibility of bias in a model) but it starts to get into tit-for-tat hand-to-hand ontology: what 'bias' means, what different kinds of bias are possible, and how even 'unbiased data' can create a model that demonstrates behaviours that colloquially and idiomatically count as biased.
But one must cut the cloth of the universe somewhere