> You just look at the data and build the best decision mechanism that can be derived from it.
How do you define "best"? That's the issue. Not all errors are equal, not all distributions of errors are equivalent even if the total is the same.
> You just look at the data and build the best decision mechanism that can be derived from it.
How do you define "best"? That's the issue. Not all errors are equal, not all distributions of errors are equivalent even if the total is the same.
That's very simple with the example in the article. Who can pay back the loan best? You just fear the answer and rather twist up your reasoning.
Even using a very short term, entirely selfish view this can be bad for the loan company. It becomes clear that blue group people are being denied loans they could well afford, and so people in that group start moving over to other providers.
If the populations of blue groups are geographically clustered, this may mean losing large portions of business in certain areas, resulting in shutting down local offices if there's a physical presence (e.g. banks).
This is also entirely aside from legal concerns.
The best model is rarely found by simply optimising a basic measure with no context.