When you are trying to solve a crime, you only know that the murderer was a nurse, it's very important to assume a valid p(murderer|gender) as well as p(gender|nurse).
On the other hand, leaking p(gender|nurse) into a candidate scoring algorithm would be a no-no.
What people seem to ask for is very interesting actually. To both learn the underlying statistics AND learn not to use it explicitly in speech at the same time. Assuming next token prediction is the learned function, these two feel a bit contradictory.