I think the weakness of this technique is in the normalization of the vectors. The close match comments don't look like mine because the content of my comments has to be massively compressed. The close matches appear to have been massively expanded.
Or to put it another way [1], cosign similarity is not enough here. Magnitude also matters here.
This is probably a case where traditional information retrieval methods should play some role. The data are not really big enough that a pure cosign similarity is warranted. [3]
[1]: a phrase that my actual Doppelganger must use. [2]
[2]: and also endnotes like these.
[3]: performative erudition is what is absent from all my matches.