Show HN: Find Your Hacker News Doppelgänger
share.streamlit.io
share.streamlit.io
I’m gonna marry one of my doppelgängers.
Relying solely on co-sign similarity, every vector is likely to be surrounded by the vectors of low karma accounts.
Or no matter which direction you travel from earth, you will almost certainly be surrounded by vacuum.
In my opinion, this service doesn't have a good S/N ratio. Could give you irrelevant information.
My accounts did not correlate, probably because they have been inactive at staggered intervals.
> We took usernames and respective comment histories from the past three years
However, putting the names in the Doppelganger search yielded very similar results and the comments of the users are from like-minded people. Well done.
From there we ran them through an embedding model and indexed the embeddings in Pinecone.
The actual similarity search is done with Pinecone. (https://www.pinecone.io)
https://news.ycombinator.com/item?id=25075318
> A reminder that BigQuery (as used in the query in this link) is the best way to play with Hacker News data; don't scrape HN data manually! The `bigquery-public-data.hacker_news.full` table appears to be up to date with the most recent HN data as well (table last updated today). However, I'm not 100% sure the query is correct for unilaterally getting all links, as running the query on the full dataset returns the same results as running it from 2006-2015. And I value my sanity enough to not fuss around with the regex.
Innit?
Or to put it another way [1], cosign similarity is not enough here. Magnitude also matters here.
This is probably a case where traditional information retrieval methods should play some role. The data are not really big enough that a pure cosign similarity is warranted. [3]
[1]: a phrase that my actual Doppelganger must use. [2]
[2]: and also endnotes like these.
[3]: performative erudition is what is absent from all my matches.
While folks are saying people they get match up with them in comment style, keep in mind due to the nature of the tool they're looking for that. Also look for comments and opinions you're not similar in to disprove the match.
I once made a Markov chain IRC bot that people would still be convinced was smart today because people discard the lines that make no sense when only looking to prove rather than also disprove
At first I thought you meant "grandparent" in the HN sense.
Given that this is my 'pharmacology' alt account, it seems the author's pretrained word embeddings still associate Coca Cola with the old recipe =)
NLP is hard!
| Username | Similarity Score |
|-----------------+------------------|
| tosh | 0.939 |
| app4soft | 0.931 |
| beefhash | 0.930 |
| joseluisq | 0.929 |
| todsacerdoti | 0.929 |
| pjmlp | 0.928 |
| rbanffy | 0.928 |
| blattimwind | 0.928 |
| formerly_proven | 0.928 |
| ducktective | 0.928 |
I identified three usernames in this table right away! tosh, todsacerdoti, and pjmlp. In fact, I like the stories posted by tosh and todsacerdoti quite often and I like the comments posted by pjmlp very often.I’m British and I’ve been living in Sweden for 7 years now, and it’s only just occurring to me that this could be affecting the way I form comments.
(I checked with correct and incorrect spellings of "dang", "pg", "TeMPOraL" and "_Microft".)
In the past I often had the impression I wouldn't get along with myself.
Now, that I'm more chill and less confrontative, I think I would like to meet myself.
Yeah, I talk a lot of shit but I've got nothing on my old self.
I'm skeptical that this tool does a good job of identifying semantic meaning of a comment, but I bet it gets the topic right.
Edit: Should be better now.
Reviewing additional posts I don't think we seem any more alike than a random pick.
Well done.
I do wonder though if the model is smart enough to correct for when you quote other people, otherwise it might be measuring who you interact with much.
BTW, I'm not sure of the privacy implications of this, maybe someone else can comment on this.
And here I was living my life thinking that Soft Cell's version of Tainted Love was the original...
https://news.ycombinator.com/threads?id=dang
and
https://news.ycombinator.com/threads?id=porphyrogene
or
https://news.ycombinator.com/threads?id=julianeon
Your site does state:
> It compares the semantic meaning of your comment history with those of all other users, and finds the top ten users whose comment histories are most similar to yours.
So maybe it's comparing comments from entire history of the account and not just the recent ones and therefore hard for me to compare? Would it be possible for you to tweak it so that it only compares lets say the most recent 30 or 50 comments?
We did this so it would work even for the less active users. But I think your suggestion might work, too.
The number of comments about the Karma being divergent on people's matches suggest you could add a simple ensemble approach with a karma heuristic on top of the vector similarity search to help results.
If the data used isn't already reflecting Karma, it could be a useful metric for representing a whole bunch of things (quality of comments, participation, length of time), and make the vector similarity more meaningful.
It would be interesting to see if users perceive the results as higher quality if you added a simple Karma similarity filter (maybe 400 points or something based on the standard deviation of the average karma score), and then returned the closest matches filtered by that metric.
Vector similarity search as a service looks like a good market. Out of interest, what would the cost translate to for running something like this in practice using the API as a customer?
We offer usage-based billing at $0.10 per GB of memory hour. (https://www.pinecone.io/pricing/) For this app, our eng team knows for sure but I think the entire index is less than 1GB so it would be just $73/month if we keep it running 24/7.
Vector similarity search is new for most companies, so we want to make it very easy to try and test stuff out in production, without cost being a barrier. Even for larger volumes (40GB+) we offer volume and pre-commitment discounts.
Also, funnily, I really like tea and I drink ~ half a gallon a day.
Corona probably got all of them :(
Edit: Nvm, just not loading in Safari
Edit: I lied, top match 'asark' tends to share in my proclivity for the run-on sentence. Topically we're not really into the same things, so it has to be something along those lines.
EDIT: those users are not even on the list of my doppelgängers.
Maybe the similarity score can be weighted by the number of comparisons somehow?
For example I suspect that if we generated 1000 random users with random gibberish comments and varied their comment numbers from 1 to 10 or so, the top similarities would be biased towards low comment-count random users. This would be because having one randomly generated comment match your style is easier than having 2 randomly generated comments match the same style.
And if that's the case then the same issue would transfer to comparing real users.
@busymom0 suggested a great solution in this thread - only do comparisons based on "n" (like last 50) comments. This way every similarity would be measured using the same number of comments and users with low comment counts would be excluded automatically.
A major reason for my first match seems to have been both accounts talking about licensing issues and specifically GPL. Other matches seem to have shared some superficially similar political perspectives on HN.
Has anyone seen a tool to do this?
Really you should just start doing it and eventually you will get there; no such thing as not clever enough.
| Username | Similarity Score |
|-----------------+------------------|
| boomboomsubban | 0.974 |
| tgsovlerkhgsel | 0.973 |
| fixermark | 0.972 |
| shadowgovt | 0.971 |
| stanferder | 0.971 |
| xvector | 0.971 |
| geofft | 0.971 |
| drdaeman | 0.971 |
| shkkmo | 0.971 |
| frombody | 0.971 |You're still wrong though :D
shantly 0.989
asark 0.988
intergalplan 0.988
jakobegger 0.988
gh-throw 0.988
moshmosh 0.988I'm wondering how matching people completely randomly under the guise of of a fake algorithm would fare. Placebo matching if you will. The results in terms of social interactions might be more fruitful.
Seriously, if you want Highlanders, this is how you get Highlanders.
That said, after reading my doppegangers comment histories, I'd totally subscribe to their newsletters.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful. Note this one: "Please don't sneer, including at the rest of the community."
Bucephalus355 0.944
freehunter 0.943
sn_master 0.943
protomyth 0.943
zapttt 0.942
cmhnn 0.942
1cvmask 0.942
sabujp 0.942
Jkvngt 0.942 (Banned)
GoinginSircles 0.942 (Throw away)I have several abandoned HN accounts because I switch to a new account after a while in order to not leave to much PII for doxxing. I don't care about karma at all.
This tool didn't offer any of them as doppelgangers, although (in theory) I should match my own style 100%.
bool AI(int id1, int id2){
return True;
}Oh gawd, now I'm sounding exactly like that guy... (I just need to add some parens.)
I’d love to see / use your work but it feels weird to participate in allowing a third party to build up an (hn-username, ip-address) database.
If you’re interested in rolling your own, a good place to start is the sentence-transformers Python package along with a KNN search service like Spotify’s Annoy.
Pinecone, of course, looks awesome as well :-)
:(
Edit: looks like it’s case sensitive, be warned people on mobile phones.
Edit2: the person I matched with the most seems a bit aggressive in their comments. I guess I can learn something from this.
journalctl 0.994
thisisweirdok 0.994
veryworried 0.993
shantly 0.993
core-questions 0.993
not_a_cop75 0.993
lambda_obrien 0.993
LeoTinnitus 0.993
xwolfi 0.993
magashna 0.993
—————————-
Edit: On a more serious note, how can we use this to find echo chambers and homogenized news sources? Keep rolling with this idea, I think it’s important.
We were hoping you wouldn't have to find out this way.
Also one was a throwaway of a guy talking about his time on shrooms /shrug
Edit: ah, one is most certainly not very nice ¯\_(ツ)_/¯
But the rule based system of law combined with the complexities of human society is something I consider interesting.
Not working for me.
> We took usernames and respective comment histories from the past three years using the Hacker News API. Then we transformed them into vector embeddings using a pre-trained model, and loaded them into Pinecone.
hardwaresofton 0.97 g82918 0.969 karmakaze 0.969 pushpop 0.969 abqexpert 0.968 breischl 0.968 forgotmypw17 0.968 dan-robertson 0.968 anaerobicover 0.968 tgbugs 0.968
definitely some good stuff to spend some time looking into though. thanks for sharing
The problem is that your matching engine assumes perfect accuracy in vectorization, but actually, the vectorization is a sample from a distribution.
The distribution of the vectorization comes from the idea that it is trained on a mix of information and randomness, and this model is just one instance of the randomness. There are some sources of randomness in the model that you control, and others in the data.
To get better recommendations, you first need to get the distributions for each element of the vectorization. The simplest approximation would be a variance around the model output. Each item should have a variance determined by the amount of information available for that sample.
Now you have another problem, which is finding the closest vector. A p-test will fail you because that is telling you the probability that you came from the same distribution, which is going to be the random distribution for almost all pairs. You might ask something like, the probability that the two points come from a distribution that is distinct from the random distribution. You’d have to form that distribution, asses probability of membership, and then probability of rejection from randomness. You could also consider doing this for each vector element and returning a negative sum of log probability to represent the amount of information shared between vectors.
But ultimately, you need a truth set to test these methods against. This is easy. You just split each person into N people by randomly assigning their comments to one of N identities. This could also be used as an element of the variance.
An easy way to start is by bootstrapping the randomness. In this case, there is probably a random initialization vector in the model. You can bootstrap the distribution of each vector element w.r.t. the IV by running the model many times with random initialization vectors. You are still using a fixed point for the training data noise, but this is a start. To bootstrap the training noise variance, you can train many models using random selection of the data, and the same IV. A good heuristic split is to decide how many times you will run the model, x, and then do random selection by 1/sqrt(x). Ex. Run 100 times with 1/10th of the data in each run. Then you have two distributions for each vector element, one for the data randomness, and one for the model randomness, but the mean from the model randomness is the most informative. Now add the data variance / sqrt(x) to get a rough approximation of the true variance.
These are all hacky ways to get a decent improvement, and the formal methods will also give great improvement on top of that, but are easy to mess up and often quite expensive to bootstrap.
Is that necessarily better? My matches felt like different people, and i like that crunchy recommendation butter more than the smooth version (unlike in real life).
Then I just mashed the button a few times and it said username found but error getting history, so I mashed the button a few more times and it eventually worked. Do probably just backend overloaded.
I skimmed through my top matches and one of them appears to have attended the same undergrad as me, so that was interesting.
Do you retain any user data? Since most users are probably looking up their own username, it would be a simple task to match and log HN usernames and their respective IPs.
Now I wonder what happened to an internet stranger.
Should be that way in general, right?
If my comments are on an island of weirdness, you may be the closest person to me while still being really far away. If your comments are relatively normal, you might have a lot of people around who are closer than me. That would make you my doppelganger, but not make me yours.
Edit: I just checked my doppelganger (barry-cotter), and I'm not even in his list :p. I've seen that username appear in a couple other comments under this post. I wonder if there are a few super normal users that a lot of people are closest to.
I have a lower similarity score with my top match, compared to my matches' similarity scores with their tenth closest matches. So it is that island scenario that you described.
FWIW, I'm not even on my #1's list.
I wish HN had a tag to identify JS-only submissions. And a filter feature to allow me to not show them at all.