Suppose the airline had a scale which worked in the following way. It will measure their weight, but add randomly generated noise of a few pounds either way. Suppose this scale is perfectly adequate for the airline's purpose - i.e. it provides a sufficiently accurate measurement to provide safety - but it no longer provides a "true" measurement.
Based on your reasoning, since this scale is no longer a perfect measurement of the true causal factor, it shouldn't be used.
In reality, passenger weight can vary between weigh-in and getting on the plane. The passenger might eat food or use the bathroom before boarding. Weight on-plane, rather than at boarding time, is the true cause of decreased safety. Weight at boarding is merely a statistical approximation to this (and admittedly a very good one).
Reality is uncertain, so it seems we are back to the same place.
Tangentially: Why stop at that level of granularity when you can easily keep slicing...As an aside, wtf is the point of FICO if not to be a measure of precisely these outcomes?
The reason for this is that using more factors is often legally problematic. Such factors are often correlated with race and regulators therefore treat using them as "redlining." For example, (income = low) AND (location == oakland) might also be predictive, but that's strongly correlated with race, and regulators will rape you if you try and use it for predictive purposes.
Also, from what I've seen, your suggesting that slicing further will eliminate the predictiveness of race doesn't seem to be true. At the very least, we often need to slice much further than our available data sets allow. For example, FICO includes a lot of data but still doesn't make race non-predictive. Similarly, in education, race is still predictive after accounting for parental income - rich blacks do worse than poor Asians: https://en.wikipedia.org/wiki/Achievement_gap_in_the_United_...
The idea that we can avoid these ethical questions by slicing data in the exact right way is probably not true. We should instead figure out our ethics so that we can do things right even if the data doesn't turn out the way we imagine.