King – Man + Woman = King?
medium.com
medium.com
But then I read the linked article [Nissim 2019] and it all became much clearer with te example in the title: "Man is to Doctor as Woman is to Doctor". If you remove "Doctor" from the results, you'll get "Nurse" instead, not because the dataset/society/... is biased, but because you inserted the bias in your model!
With this small "optimization" (and others presented in the article, such as the "threshold-method" and hand-picked results from the Top-N words), it's trivial to use the anology method to show that any dataset is biased, since you are filtering the unbiased results!
One a side note, I wish articles like this were more popular. I find that it's really easy to use AI techniques "the wrong way" and there's are a lack of articles pointing to common pitfalls (and, to make things worse, there are a lot of blog posts doing things wrong, which only validates bad methodology).
[Nissim 2019]: https://arxiv.org/pdf/1905.09866.pdf
Thus—given that we’re defining words based on their centroids of usage in a historical corpus—if a woman is a monarch of a kingdom, “king” is a tighter historical fit to describe her role than “queen” is.
I did the same thing with Japanese and Western movies, and the result was not due to the movies just being close. See https://strandmark.wordpress.com/2017/07/08/the-vector-from-...
But I think you are absolutely correct that very few articles explicitly acknowledge and handle these types of common pitfalls. Especially more intro-level articles. Personally, I'd love to read more about effective strategies for ensembling different algorithms to work around some of those pitfalls. For example, the OP points out that Word2Vec underperforms on lexical semantics. I wonder if you could use a different algorithm that performs well on lexical semantics but poorly on other categories as a supplement, with some third algorithm deciding which to use in any given case.
Your last point sounds like a cool idea! Using those more in-depth metrics to find weaknesses and see if other, complementary algorithms can fill the gap.
I think it is also important to note that this is an argument over which of two methodologies is best. The answer likely is that it depends on the use case. Researchers are not being accused of using underhanded methods, manipulating data or cheating.
And bias can certainly be of different magnitude. A weighted dice can have a bias that shows up after only 10 rolls and another one can only show it's bias after 1000.
A random sample is not a product of a biased sampling method, but it can be (and almost always is) biased, it's just if you take enough such samples, the biases will tend to offset, which is why we say the sampling method isn't biased.
I think that this hypothetical example would be evidence of bias. "woman:doctor :: man:nurse" as an analogy only works because of gendered implications of 'doctor' versus 'nurse', so if that were a robust output of the model it would imply that the word2vec found that gender was of great significance in the manifold near 'doctor'. (The analogy would just be applying a negative weight to the gender basis, but that's mathematically fine even if not customary.)
In a mathematical context, if you have something like this: 2:0::4: (two is to zero as four is to what?) then acceptable answers might be 2 (rule is: subtract 2), or maybe it could be 1 (take half, subtract one). But it couldn't be zero. (rule can't be "multiply by 0").
I could be wrong but I believe that's how it works.
Particular tests may adopt such a rule, but this doesn't apply to analogies in general. Particularly, consider the case where Alice and Carol and both daughters of Bob: Alice:Bob::Carol:Bob is a perfectly valid analogy; a:b::c:b is commonly phrased along the lines of “a and c are similarly situated with respect to b”.
NLP-wise, is there "a right way"?
Actually, some years ago I bumped into a similar problem to the one discussed, where someone wanted to use NLP to show gender bias in a dataset, and hit one of the common pitfalls that I mention (I hope that I don't start a flamewar by sharing this story):
Here's what they tried to do: 1. Fetch a list of ~800 atendee names from a Portuguese tech conference (it had an official API with user profiles) 2. Download a dataset most common male/female names for newborn babies in Portugal and America for the latest 3 years 3. Train a naive bayes model on the downloaded dataset and use it to classify the antedees into male/female
After doing that, the algorithm returned something like "8 female attendees and 792 male antendees".
I found this particularly strange (considering that I knew more than 8 women that attended on previous years), so I took a peek at the antendee dataset and found that: - There were some users using a fake name (including one organization account) - There were certainly more than 8 female antendees, and at least 6 were named "Inês" (female name)
After discussing this with the ones involved, we found the problem! - The dataset was not being normalized (it was being trained with "Ines" and tested with "Inês") - The naive bayes[1] implementation used, when faced with a completly new input, outputed the most common class of the training dataset
In the end, the final result was closer to "80 female, 705 male, 15 unknown", which is a much more believable result (closer to the typical distribution of Software Engineer students in Portugal).
Note that the author wasn't trying to deceive anyone, he just tripped on some common pitfalls (forgot to normalize the data and used an off-the-shelf implementation without looking into the details).
[1] There was only one attribute, so implementing this without using a naive bayes library was actually easier and produced the correct results.
Technically in England Queen is an ambiguous term, queen regnant is the actual ruler, queen dowager is the widow of a king etc. But, again several European an other female rulers used the male term.
Hatshepsut crowned herself Pharaoh and maintained an elaborate legal fiction of maleness, because ya know she was in charge and could do what she wanted.
Hungary had two female kings Mary of the House of Anjou, and Maria Theresa.
Etc etc.
So if people actually do the calculation King-Man+Woman and it comes closest to King, than they should report "King-Man+Woman~=King" and not "King-Man+Woman=Queen" (only because that's what they expected).
The problems materialize when you're just cherry picking for results.
The query would give you a ranked list of the closest word vectors with scores that indicate how good the match is.
But, what are you really doing there? Well, you are operating with vectors. And what do those vectors represent? Not the semantics about the words themselves. The vectors encode the context in which they appear. They encode what other words tend to surround them.
And there are more limitations. The most obvious one is polysemy. But even without polysemy, different words might have more or less "diffuse" representations. The more technical, precise and uncommon a term is, the less "diffuse" its vector is. King and queen, man and woman are far from precise, technical or uncommon terms. And therefore, there's a lot of "diffusion". Again, the problem here is our expectations versus the results we see in practice. But the models are what they are. And they are not magic.
Also, if we think about it in terms of decision manifolds, it seems the distance between queen and king is too large for the simple - man + woman to have an effect. Why not scale that substraction, so it leads to a change in predicted class without removing king? But of course finding a justifiable weight would be hard..
[1] http://bionlp-www.utu.fi/wv_demo/ (making sure to select the English model)
In fact, the plural isn't removed. You can see the effect by analogizing A:B :: A:?.
For man:king :: man:?, you get [kings, queen, monarch, crown_prince] as the top 4. For man:king :: woman:?, the results are [queen, monarch, princess, crown_prince], with 'kings' as #6.
Of course, your model may vary.
† Because it doesn't measure performance at something people typically want to do with these models in real life - there's a strong "party trick" component to the word analogy stuff.
My own sense is, the real problem that the space needs to contend with isn't that the math isn't elaborate enough. It's that we're still waiting for someone to come up with a really good way to deal with polysemy, and with multi-word phrases that form a single semantic unit.
Yes, there have been. https://github.com/lmcinnes/umap_paper_notebooks The author of UMAP shows impressive results with 3 million word vectors
For dealing with multiple meanings, It's already been done. Sense2vec resolves most of these issues and a wordnet integrated version of word2vec or the newer "XLNet" would be state of the art by a long shot but no one seems to want to implement it so the world waits longer for good NLP models I guess...
It relies on word sense disambiguation, which tends to be one of those very language-specific things, and so I'd expect (but haven't verified) that, like other techniques that rely on language-specific bits, it wouldn't work as well on most non-English text. And the most interesting polysemy problems aren't to do with part of speech. They're things like "apple-as-in-food" vs "apple-as-in-computer", or figuring out that "The Big Apple" doesn't have anything to do with either of those. What would be really interesting is dealing well with jargon, slang, and terms of art.
As far as those notebooks, is there one in particular I should be looking at? I might have missed something, but the stuff I saw basically just demonstrated, "Hey, we can handle a lot of training data really fast." What I'd be more interested in seeing is, "Hey, plug us into your document classification pipeline and your performance (as in accuracy) metrics won't know what hit them."
edit: For a more concrete example of what I'd like to see, and going back to the analogy task: The holy grail I'm looking for isn't "king - man + woman = queen". It's more like "software engineer = programmer", and also "software engineer != software + engineer".
Oh and you can use UMAP to concat tons of vector models together and all other side data for super-loaded embeddings
It's been done: contextual embeddings - ELMO, BERT.
In 3d space for example (much simplified from 300d), things that are separated along a vertical axis are usually very different in kind than things separated even by a great distance on the horizontal axis (100m below you is likely to be much different from 100m in front of you).
Or imagine a sky scraper most of the "sameness" lies in a small horizontal region (one block say), but a vast vertical region (100 stories).
This is a simplified analogy to word vectors, but the point is that because two words are "close" in 300d space, if we don't understand what those dimensions mean, we can't say which one is more likely to be "similar" for a specific pair of words (King/Queen vs Mouse/Cat). Using euclidean or cosine similarity may or may not be relevant for one particular case.
Besides that it's actually "King - Man + Woman = Queen [Regnant]", which broken down into smaller pieces should be "King - Man = Monarch", "Monarch + Woman = Queen [Regnant]". Then you also have the complicating factor of "Monarch + Wife = Queen [Consort]" and "Monarch + Husband = Prince [Consort]". It seems obvious from this, and "Husband - Man = Spouse" and "Wife - Woman = Spouse" that "Monarch + Spouse = Consort".
This allows a little thought experiment. If the king marries a man, you have a king and a prince. That's fine; no ambiguity there. If the queen [regnant] marries a woman, you have a queen and a queen. This is confusingly ambiguous. Do you rename the queen [regnant] to king, or do you rename the queen [consort] to princess [consort]?
So the assumption is that words from similar context should be similar. But you're always going to miss out on some words that are similar but do not appear in similar contexts.
The words might not appear in a directly shared context, but given their semantic similarity, they should share more contexts of distance 1 (or something along those lines) than an arbitrary pair of words.
King - tomato + potato = Queen
?
(but sure one could also pick queen, prince, royal form the list...)
Just tested it here: http://vectors.nlpl.eu/explore/embeddings/en/calculator/#
And it gave me 0.63 King, 0.6 Prince etc...
is kinda sexist.
"King – (Man + Woman)"
and
"(King – Man) + Woman"
Must have different outcome? I do. No wonder Word2Vec sucks! Words are not elements of vector space :)
5 - (2 + 3) =/= (5 - 2) + 3
The same holds for elements of a vector space.
I was about to say you were wrong, but you are correct, and it is a bit unintuitive why - https://www.quora.com/Is-vector-subtraction-associative
It's part of the learning curve, and I spent quite some time to learn the basics
In the "King" example, you're adding and subtracting two words that are probably very close already, so if you want to find "something else" besides itself, you need to exclude it. For some problems it might make sense, for some others it might not.
That's not an accurate description what Bolukbasi et al (2016) [0] did. In particular, they do not list x close to lovely + he - she and then pick arbitrarily from that list. Instead, they explicitly reject that approach (see appendix A), because they're looking for pairs of words that are maximally gendered. They do that by finding x and y such that the angle between x - y and she - he is minimized. Since the task they're solving is different, you can't fault them for getting different results.
Does not look very scientific to me.
Did you even read the article? Its whole point is that sensationalist conclusions like yours are wrong.