Those aren't mutually exclusive. Technically speaking, the model can return results "almost at random" and still demonstrate a bias against any particular attribute if that bias is evident in the underlying training dataset.
If there are strictly fewer women in the underlying training set, the model can still return something resembling a uniform distribution of candidates while exacerbating the diminished representation of women.
To give a concrete example: you have a bag of blue dice and red dice. There is a supermajority of blue dice in the bag. Your algorithm selects a single die out of the bag on every iteration. The output sequence of dice numbers appears uniform, but there are more blue dice than red dice in the output sequence.