It really isn't ambiguous. What is less clear is what to do about it, and where that even makes sense.
There are few meanings of bias that are important to keep clear about when discussing things. (1) There is the statistical sense of a biased estimator - one with a consistent trend in it's error. (2) There is the notion of bias introduced by data sampling. No matter how perfect your algorithm is, if the training data is a poor sampling of the general population you are targeting, you are likely introducing systematic bias (for example, early face detection approaches had best performance on Caucasian, male, college aged faces - people were using the data it was easiest for them to collect[1]. Finally (3) if the data you have access to has encoded a systematic bias, even if the first 2 have been avoided you are at best able to reinforce that bias.
Here we are mostly talking about latter one, and problems being encountered it (and a bit of 2). This is exacerbated by a combination of machine-learning and AI people being fairly unsophisticated about data on average (as opposed to data handling), and popular techniques these days (specifically, deep learning) de-emphasizing feature design making it harder sometimes to see what is happening.
Nobody serious I have seen is advocating engineering outcomes into AI to adjust outcomes.
For the sake of argument let's assume there is in fact an over representation of men (compared to women) in technology relative to capability and desire. And that we are designing an AI to make or aid hiring decisions for entry level jobs. The answer then is not to engineer in a quota for women applicants, but to remove gender entirely from the training and evaluation inputs. This would give you exactly the desired outcome, no?
You refer to the type (2) bias issue if we systematically under represent women in tech in this case, which would be a problem. However the article is focused on the type (3), which is not merely an issue of what is in the data set, but what you are trying to do with it.
The deep issue here is that deep learning approach of throw everything at the inputs and let the network sort it out will capture both (2) and (3) types of biases whether or not we are aware of them. At least in areas where there is very objective proof of potential problems in the historical data (e.g. redlining impact on mortgage decisions) we could get ahead of it and normalize the inputs. But what about areas where it is less clear cut or more contentious?
There is the fact that humans fail in exactly the same way. If I'm making the decision to hire, it's very difficult for me to avoid my own biases. I think one of the things people are concerned about with applying ML techniques to things like this is it gives the appearance of being more likely to be unbiased, and potentially provide cover from those who would like to benefit from it.
I suspect the real answer is to start thinking about these things the same way we have learned to thing about security and cryptography. In other words, for serious work a system is not considered ready for prime time until real professionals have tried to find the weak points and break it.
It is also worth noting that this is not necessarily a problem so much as an operational feature of the approaches one should be aware of when designing and using them. And in some applications it probably a very significant problem.
[1] one of the interesting thing about this and similar problems is how completely predictable it was, and how surprising it was for practitioners to discover it. In this way the ML community needlessly recapitulated lessons learned by other disciplines decades earlier.