>If you start a query with “Mary Lee Pfeiffer”, you’re not going to get very far because neural networks aren’t equidistant grids of points (besides the fact that she may not appear very often under that version of her name.) They’re networks of nodes, some with many connections, some with few. One of the ways you optimize large models is by pruning off weakly connected regions. This may come at the expense of destroying B is A relationships for weakly represented entities.
Am I missing something here? This paragraph reads like complete gibberish to me.
Also, I don't buy the experiment at the end. If you fine-tune the model with exclusively Tom Cruise data, I want to see proof that it doesn't just answer "Tom Cruise" all the time. I want to see that it says Tom Cruise wrote Aces in the Stream, but doesn't say Tom Cruise wrote Kings in the River.