So, the issue is more subtle than that - if it was just a matter of including demographic data, there wouldn't be an issue. The problem is that the "B" column is _not_ measuring how integrated the neighborhood is- it is a transformed value whose calculation begins with data about integration levels, but the details of how and why that transformation is performed are super important. The calculation the original authors did to produce that column rests on a model about the relationship between a neighborhood's level of integration and its property values, and that model's assumptions are frankly racist and also factually incorrect (and were known to be incorrect at the time of its original publication back in the '70s). As a result, if one were to _use_ the "B" column in a model, one would be getting results that at best were wrong and useless, and at worst would make a model that literally encodes broken 70s-era ideas about how real estate and race interact in the US.
And the transformation itself is non-invertible, so it's not possible to recover the original values for about 7-8% of the rows in the dataset. The commit diff links to a thorough investigation of the data[1] in which the author takes a crack at linking up the ambiguous rows with the original 1970 Census data that supposedly went in to generating this dataset, and long story short, it looks like the original dataset's authors may have made some errors in their calculations on top of everything else.
1: https://medium.com/@docintangible/racist-data-destruction-11...