And if those results change, so will the algorithm's outputs! But asking the algorithm to
make the change seems to be a bit much.
Honestly, my preferred method of solving this would be to train the algorithm on a data set with all of the forbidden values included along with anything else the creator feels relevant - zip code, income, familial status, favorite sport, education - and then when running in production, against real people, don't give it the restricted information. Yes, you could theoretically extract race, gender and other protected stats from the information the algorithm actually uses in prod - but it has no incentive to, since a less-noisy signal is already provided.
For instance, suppose the optimal algorithm for your data set is some linear function of X, Y and Z - let's say X+Y+Z to keep things simple. X,Y and Z are all normally distributed variables, mean of 0 and the same standard deviation. Y has a 0.5 correlation with X, and a -0.5 correlation with Z. If not provided Y, your algorithm might come up with 1.5X+0.5Z as an approximation - extracting a bit of the signal for Y from the things it does have access to. It's suboptimal, but better than just X+Z. Unfortunately, Y is verboten - we're not allowed to discriminate on it, and this approximation ends up with results that track Y. So instead we train with X, Y and Z as inputs, so the derived model is X+Y+Z - and we can drop Y from that model in production, leading to a model that (while less accurate) shouldn't unfairly track Y.