Show HN: Using stylometry to find HN users with alternate accounts
stylometry.net
Here's Paul Graham:
https://stylometry.net/user?username=pg
Here are some frequent HN commenters: (EDIT: Removed due to privacy concerns)
stylometry.net
Here's Paul Graham:
https://stylometry.net/user?username=pg
Here are some frequent HN commenters: (EDIT: Removed due to privacy concerns)
The most interesting thing is that my writing style changed pretty drastically since a decade ago. Searching for my oldest account matches my earliest usernames, whereas searching this account matched the rest.
The details of the algorithm are fascinating: https://stylometry.net/about Mostly because of how simple it is. I assumed it would measure word embeddings against a trained ML model, but nothing so fancy.
I put in my username and found my pre-echelon alt, possibilistic.
(Echelon was taken when I registered possibilistic, but it must have been unused and dropped.)
As far as I’m concerned, it’s the killer feature of the app. The top 20 results may be noisy, but the bolded results have a signal to noise ratio close to infinity.
But I don't think people should be making the assumption that bolded results are definite alts, which sillysaurus' comment reads like.
It’s one of those ideas that makes the tool substantially more effective, yet never would’ve occurred to me. It’s like the simplicity of pg’s “a plan for spam” algorithm: deceptively simple, but (like scrubbing dishes with fingers) works really well.
That is absolutely all this will be used for. This is a dangerous tool that serves no real world purpose.
None of them are me (and you were the only one I recognised and thought "yeah, I can see where it gets it from"...)
Have you considered doing rune rather than word ngrams? I can imagine that might be prohibitively expensive, but I really don’t know. I did something like that long long ago in C for automatic document language detection. It was quite accurate.
> sillysaurus2
Tbf a human could have found a bunch of them relatively easily
https://news.ycombinator.com/item?id=17944293
The approach I took was a bit different, but also no ML required.
The real trick is pruning and going cross platform. There are around 100k active HN accounts (meaning posts a few times a year), maybe 200k if you count at least one post a year. But <10k that post weekly.
It’s a very small space to try to compare so simple methods will work fine.
I create new accounts on a semi-regular basis because I think cliques are the most corrosive factor to social media. Any time my account gathers enough upvotes enough I destroy it for another.
I had four accounts. None are over 50% confidence, but when I look at any one account the others are consistently #2, #3, and #4.
Now I’m thinking very carefully about what words I use to avoid linking this as the 5th account.
Also, the cosine of the vectors of word frequencies conflates author-specific vocabulary and topics; in other words, my account is grouped (with >51% similarity, according to the demo) with someone probably because we wrote about similar things. A strong stylometric matcher ought to be robust against topic shifts (our personal writing style is what stays constant when we move from writing about one topic to writing about another topic, just like our personality is what stays constant about our behavior over time - of course styles do change, but the premise then has to be that such changes happen very slowly).
Stylometrics/authorship identification is interesting and has led to some surprising findings, e.g. in forensic linguistics (Malcolm Coulthard wrote several good books about the topic).
This paper lists some other features that could be used and compares a bunch of techniques: https://research.ijcaonline.org/volume86/number12/pxc3893384...
Interesting. I was expecting to be grouped with other Russian speakers and I am (based on some nicknames). But I thought the most telling feature will be exactly word order - it’s absolutely relaxed in Russian. Word frequencies? Well, probably the absence of articles, lol (but I swear to God that I often spend some extra time trying to insert as many articles in my texts as I could).
”Language consists of sentence constructs, choice of words, and expression of style. Accordingly, an idiolect is an individual's personal use of these facets. Every person has a unique idiolect influenced by their language, socioeconomic status, and geographical location.”
This is a fascinating way to find similar HN users who aren’t the same person. It’s a surprisingly great recommendation engine. “If you like pg, you might also like…”
Sure, the privacy concerns are valid, but the cat’s out of the boot. Might as well enjoy the benefits.
montrose is almost definitely pg. Someone who talks about ancient history, Occam’s razor, VCs and startups, uses the phrase “YC cos” (relatively uncommon), etc. https://news.ycombinator.com/item?id=17112567
Nicely done. One of the best hacks I’ve seen in a long time.
I had this hunch too. It's either pg or someone trying really hard to be pg.
> someone trying really hard to be pg
describes half the site.
I think these are all common topics among HN readers and commenters.
It's my first time hearing that variant. Usually its, "the cat's out of the bag" where I'm from.
Do you mean boot in the UK sense, what Americans would call the trunk of a car? Or do you mean a sturdy piece of footwear?
Obligatory xkcd https://xkcd.com/2390/
It’s a fun game, too. I wish I’d used “the cat’s out of the hat,” but I didn’t think of it till later.
I vaguely remember one of the metaphors in the essay was about a chicken coop melting, or something like that. It was vivid enough to leave a big impression.
“ Dying metaphors. A newly invented metaphor assists thought by evoking a visual image, while on the other hand a metaphor which is technically ‘dead’ (e. g. iron resolution) has in effect reverted to being an ordinary word and can generally be used without loss of vividness. But in between these two classes there is a huge dump of worn-out metaphors which have lost all evocative power and are merely used because they save people the trouble of inventing phrases for themselves.”
(It’s remarkable how often a vague description can yield an HN comment with an answer from a clever sleuth like yourself. Much appreciated.)
The 2nd example also loosely falls under the classification of malaphor.
https://thehabit.co/knowledge-is-power-france-is-bacon/
> When I was young my father said to me: “Knowledge is power, Francis Bacon.” I understood it as “Knowledge is power, France is bacon.”
> For more than a decade I wondered over the meaning of the second part and what was the surreal linkage between the two. If I said the quote to someone, “Knowledge is power, France is Bacon,” they nodded knowingly. Or someone might say, “Knowledge is power” and I’d finish the quote “France is bacon,” and they wouldn’t look at me like I’d said something very odd, but thoughtfully agree. I did ask a teacher what did “Knowledge is power, France is bacon” mean and got a full 10-minute explanation of the “knowledge is power” bit but nothing on “France is bacon.” When I prompted further explanation by saying “France is bacon?” in a questioning tone, I just got a “yes.” At 12 I didn’t have the confidence to press it further. I just accepted it as something I’d never understand.
> It wasn’t until years later I saw it written down that the penny dropped.
Not necessarily, you might be thinking of malapropisms but yes probably a closer word would be the general term: protologism.
Another commenter added some useful info on the evocative alteration of metaphors [2]
Easy to overuse then people just get annoyed though…kind of like commas, I suppose.
- Is bolded on pg's page
- Mentions yoga
- Talks about Lisp often
- Talks about YC often
- Talks about kids
- Links to Paul Graham's website
- Says he uses vi
- Writes exactly like you would expect pg to write
Is that what PG would say?
Users freq. say they will pay for something but back down against other things.
https://news.ycombinator.com/item?id=16785542
https://twitter.com/paulg/status/1362369484036653058
I think you are very likely to be wrong.
I think, the similiarity has to be in the high .80's to suspect that it's the same individual.
Excerpts from wiki:
> Before the publication of Industrial Society and Its Future, Kaczynski's brother, David, was encouraged by his wife to follow up on suspicions that Ted was the Unabomber.[91] David was dismissive at first, but he took the likelihood more seriously after reading the manifesto a week after it was published in September 1995. He searched through old family papers and found letters dating to the 1970s that Ted had sent to newspapers to protest the abuses of technology using phrasing similar to that in the manifesto.[92]
> In early 1996, an investigator working with Bisceglie contacted former FBI hostage negotiator and criminal profiler Clinton R. Van Zandt. Bisceglie asked him to compare the manifesto to typewritten copies of handwritten letters David had received from his brother. Van Zandt's initial analysis determined that there was better than a 60 percent chance that the same person had written the manifesto, which had been in public circulation for half a year. Van Zandt's second analytical team determined a higher likelihood. He recommended Bisceglie's client contact the FBI immediately.[96]
> In February 1996, Bisceglie gave a copy of the 1971 essay written by Ted Kaczynski to Molly Flynn at the FBI.[87] She forwarded the essay to the San Francisco-based task force. FBI profiler James R. Fitzgerald[98][99] recognized similarities in the writings using linguistic analysis and determined that the author of the essays and the manifesto was almost certainly the same person. Combined with facts gleaned from the bombings and Kaczynski's life, the analysis provided the basis for an affidavit signed by Terry Turchie, the head of the entire investigation, in support of the application for a search warrant.[87]
Also interesting:
> Ross Ulbricht aka Dread Pirate Roberts, the mastermind behind the infamous Silk Road site which served as a black market for drugs, weapons and fake documents was also well aware of the potential danger of stylometry being used against him. At the time of his arrest in a San Francisco public library, the FBI captured images of his laptop screen as evidence. Guess what what he had bookmarked — “Science of Stylometry.”
https://medium.com/svilenk/the-case-for-anonymity-12db114f0c...
I often wonder if stylometry can be used to positively identify a person based not on general word frequency, but by a single phrase or two which are rare in general but commonly used by the individual. In theory this could be relatively easy to find given a large corpus. You'd pick out the top few n-grams for short phrases by an individual and identify the ones which are most overly-represented compared to the rest of the population.
I appear to be a well-educated, over-confident know-it-all.
Don't we all?
> In the summer of 1997, I was 28 years old, and I decided that after years of thinking about writing a novel, I was simply going to go ahead and write one. There were two motivations for doing so. First, I was simply curious if I could; I'd had up to that time a reasonably successful life as a writer, but I'd never written anything longer than ten pages in my life outside of a classroom setting. Two, my ten-year high school reunion was coming up, and I wanted to be able to say I'd finished a novel just in case anyone asked (they didn't, the bastards).
> In sitting down to write the novel, I decided to make it easy on myself. I decided first that I wasn't going to try to write something near and dear to my heart, just a fun story. That way, if I screwed it up (which was a real possibility), it wasn't like I was screwing up the One Story That Mattered To Me. I decided also that the goal of writing the novel was the actual writing of it -- not the selling of it, which is usually the goal of a novelist. I didn't want to worry about whether it was good enough to sell; I just wanted to have the experience of writing a story over the length of a novel, and see what I thought about it. Not every writer is a novelist; I wanted to see if I was.
The most noticeable similarity is that we both clearly have strong opinions about some things, and like to share information, but also like to be clear about our unknowns or opinions. So, lots of “sounds likes,” “probably,” “could be” and so on.
The downside is, I guess, this could be seen as a bit weasel-word-y or indirect.
Commonly called just “hedging” like hedging your bets.
I do think it is an under-emphasized aspect of honesty, though, that we should be clear about our level of experience/understanding. Especially online — people like to discuss things, even (especially?) when we are just getting started. So if we’ve picked up opinions through osmosis and we start repeating them without testing them, we’re really just amplifying some possibly-incorrect viewpoint (and if we’ve picked it up, there’s a good chance it is already widespread in the community, which is bad if it is wrong).
And I mean, more concretely a measurement is not complete without the error bars!
Often this doesn’t really matter, because it is just chit-chat anyway. But it is nice to keep in mind.
there are many languages that encode this info as mandatory grammatical affixes, it's called evidentiality.
I find it interesting that the first example they use in the Wikipedia article is Turkish. I’ve only met a couple Turks, but they were all quite good engineers. I wonder to what extent embedding this kind of information in the language helps organize your thoughts.
This was certainly eye-opening.
Update: It’s actually a little strange that reading through some of the matches it’s not just style that overlaps but perspectives in quite a few cases too. I’m definitely not the unique little snowflake that some others are finding themselves to be.
I’m pretty sure participation in HN is a 99% sure filter for being called this many times in one’s life.
And yeah, there's a bunch of high confidence (.6-.8) hits for that account, and from a quick browse of the comments of the recently active ones, they look really likely to be alts. Like, all three that I looked at had comments that made it very clear it was this person writing pseudonymously. (E.g. writing on their signature issue, and saying they couldn't go into more detail due to fear of self-doxxing; or somebody literally saying that the alt's claims reminded them of the public writings of the notorious guy years ago).
Obviously I'm not naming the account, but this functionality turned out way creepier than I thought the moment I tried it on the account of somebody who has a reason to disassociate from an existing public persona, but still wants to participate here.
(I don’t know who the comment is talking about, which is how it should be. There’s no need to blow someone’s cover in a highly visible way. Even if they were satan, they’d still be welcome on HN as long as they’re writing substantive, interesting comments that follow the guidelines.)
I would say it's obvious why one might respect that wish (do unto others...), but I'm also aware that my and my culture's sense of privacy goes further than many others'.
Hopefully this raised awareness means that people who actually need anonymity will be more likely to know to take precautions.
I've been thinking about this a bit, and I've landed in that having a stable identifier across ALL comments & posts is a poor default. We still probably want some coherence, at minimum within a thread, eg to follow a back-and-forth. The site itself may also use stable identifier for abuse prevention. But there's no reason one should have the same username externally traceable for posts about completely different topics.
In practice, this could be done with low friction pseudonym creation, which all ties to the same account privately.
It is probable that persons of same cultural origin will have similar writing style and vocabulary. It is also probable that persons of same cultural origin would have same relationships with the world as a whole, they would like same things and dislike other same things.
So, in my opinion, it is possible that you have found not only alternate accounts (score above 0.7), but accounts of people with same cultural origin (ones that are around 0.6).
I think your tool should have internal embeddings for each of the user. Also, most probably your tool uses cosine similarity for a search.
Thus, I would like to suggest a feature: recognize simple arithmetic operations over user's embeddings, such as "thesz - 2 * patio11". It will make things even more fun, this way we can find users who are like me and much not like patio11. Even simple additions and subtractions would suffice.
(an idea is taken from properties of word2vec embeddings)
Your tool is thought provoking. What I discovered with it made me think about my use of language and what other languages (body, imagery, etc) I use differently because of who I am. Which made me think about my favorite underrated superhero Cypher [1] - would his innate ability to understand languages make him best detective ever?
[1] https://en.wikipedia.org/wiki/Cypher_(Marvel_Comics)
Thank you!
Now they know just who cares about which alternate accounts. They know!
They freaking know, man!
You have all fallen for their ploy. Fools!
What I found was worth visiting the site. Somehow notably many accounts with (relatively) high similarity to mine's are sharing at least one of my personal traits.
Which is fascinating, to me.
And I think is worth to be noticed by others - what and how you write can disclose who you are.
(Or does it?)
On an unrelated topic, I'm starting a service to write comments in the style of others to provide plausible deniability for other alt accounts. Rates negotiable.
No other hits above 0.5, so I guess that either makes me pretty unique as a commentator or my English is broken in a unique way.
Found this, it includes 250k usernames, but it's not there. https://www.kaggle.com/datasets/hacker-news/hacker-news-corp...
https://console.cloud.google.com/marketplace/details/y-combi...
A funny thought — my “matches” cap out at around .56. Having false positives* in a tool like this might feel like a “bad result” but actually I think it just means that if someone were running this sort of tool across the whole internet, I’d be relatively easy to correlate, while your identity would be intermingled with your .6-.7 partners.
*actually they aren’t really even false positives because the tool doesn’t promise to detect alts in the first place, just find similar styles.
Hmm isn't a spot check of comments somewhat tautological, since that is how the tool identifies alts (rather than something like IP address or time of day)? If this had been promoted as "find accounts with similar writing style to yours" would people immediately assume alts?
If some accounts are found to be stylometrically similar, and then a visual inspection also shows them all stating similar opinions, that latter piece of data is a strong signal.
I don't think the list of pg alternate account is accurate. I checked a few. They have many oneliners that is typical of pg, but the topics and style don't look similar.
I searched a few more and got better results. :)
I searched myself (that I know that I have no alternate accounts). I recognize a few users that are interested in similar topics, and I discuss/upvote them many times. But I didn't recognize most of the user of the list.
It's based purely off frequency of the 200 most common English 1 word phrases, 2 word phrases, 3 word phrases, 1 character sequences, 2 character sequences, and 3 character sequences. Topic does not really have anything to do with it. If I had more time I probably would've done a smarter model that accounted for things like that.
Another is form Argentina, so I guess the native language leaks, for example using words derived from latin that are not idiomatic.
And there are a few more, that is a honor to be "confused" with, but I have no clue why.
I don't do throwaway. I either post or STFU. I also STFU on darknet. Its why I found it fun to read/lurk on things like I2P back when it was new. And I know that on a pseudonymous account it is only a matter of time until it can be linked to another pseudonymous account. It would not surprise me if stylometry was used on Dread Pirate Roberts or the people behind The Pirate Bay or the people behind Wikileaks (Assange's sockpuppet accounts). Such can also have been used to verify afterwards instead of beforehand. Though with TPB since it was on clearweb an advanced adversary could have used correlation/timing attack to figure who wrote what.
I'm having fun times recognizing other Dutch people though their usage of English language. For example, a distinctive word I see Dutch people use a lot is 'oke' instead of 'OK' or 'okay'. Its a red flag the person is native Dutch. I wonder if there are stylometry tools available for figuring if someone used physical vs touchscreen keyboard (I used Glider to write this post, spellchecker unavailable).
And yes, organizations like secret service and police should use such tools as well. It is a known tool, why not use it for good? As with any tool, it can be used for good and evil. On HN this could be useful for the mod team (AFAIK nowadays only dang) to find banned people's sockpuppets. Cross-community could also be a fun project: find a HN user's Twitter or Reddit account. And I hope this method is also used to find Russian trolls on social media.
I wonder if this explains our similarity. And if so, could we tweak the algo by e.g. Removing text that is prepended with ”>”
Also, since this uses only word frequency, there are probably relatively easy improvements to make that would make it even more powerful, like looking at particular runs of words that are unique. Some expressions or figurative language only show up in combinations of words, and tend to be highly style specific.
Sometimes people insist that's all role-play and irony; others insist that if it ever was, it certainly isn't now.
But regardless, I remember pre-2005, and it wasn't all like what I saw the two times I looked at 4chan. Bits were. Bits were much worse. But mostly, mostly, people were kinder… at least, unless political tribalism came up.
I'm 0.566 correlated with logfromblammo -- and while we are definitely not the same person, I could easily imagine writing a sentence such as:
"For some bizarre reason, management has not yet assigned a task to their programmer underlings to automated themselves out of existence. I can't imagine why."
which is theirs, not mine, from about a year ago. I like that.
On the other hand, I'm nearly as correlated with peterwwillis: 0.5485 -- who has no comments and no submissions.
This is due to the Firebase API not updating when users ask the admins to move their comments to another account.
As you're typing out a comment the software gives you a list of accounts you're becoming similar to. That way you can adjust your writing as you type.
For example, as somewhat illustrated here, your personal vocabulary is a kind of fingerprint. As you mention, using a thesaurus can somewhat alleviate that, but if a thesaurus is only changing a small % of your words, then it will only have a suitably small % effect upon analysis.
To go yet further might (I suspect!) entail methods such as directly lifting and using other people's sentences to convey your own thoughts. But even then, "your own thought patterns" are still informing the manner of the post, to some extent, so over time increasingly robust analysis may still find patterns to hook into.
They were always communicating in some kind of meme-russian, and their texts were funny to read. [1]
I believe their writing mostly defeated this kind of analysis, at the cost of looking like idiots (which was probably the reason no one sent them crypto-dollars to buy that stuff exclusively).
Here's an excerpt:
"Attention government sponsors of cyber warfare and those who profit from it !!!!
How much you pay for enemies cyber weapons? Not malware you find in networks. Both sides, RAT + LP, full state sponsor tool set? We find cyber weapons made by creators of stuxnet, duqu, flame. Kaspersky calls Equation Group. We follow Equation Group traffic. We find Equation Group source range. We hack Equation Group. We find many many Equation Group cyber weapons. You see pictures. We give you some Equation Group files free, you see. This is good proof no? You enjoy!!! You break many things. You find many intrusions. You write many words. But not all, we are auction the best files."
[1] https://archive.ph/20160815133924/http://pastebin.com/NDTU5k...
(FWIW, it didn’t find my throwaways; my own model didn’t, either, because I knew that word choice wasn’t enough to avoid being outed by stylometry)
Edit: by bigrams and trigrams, I mean reducing word to their parts of speech labels and using THOSE as word tokens. You’ll find that native English speakers have higher weights on some phrase construction patterns than, say, folks from Romania. TF-IDF is useful for these POS-grams (just made that word up) as well.
That is a very good idea and when I update the site that will almost certainly be included :) Any other tips? Been reading papers for ideas and I think I may have to ditch the cosine similarity and go for something fancier soon. Thank you
“Find hot single women who write just like you”
In my experience I don't see a relevant list of potential matches aside from gender and age preference, it's all completely random, even frequently I see people outside the settings I've specified (i.e. men or older women).
It's so easy for something like this to be turned into a tool for a witch hunt, targeting innocents.
Will be interesting if we could plot the writing style divergence over time.
https://stylometry.net/user?username=stavros
The next person is 30% less certain, that's huge! This would basically identify any alt I might have with near certainty.
https://stylometry.net/user?username=rogual
I'd have thought this stylometry thing would be commutative.
0.1, 0.2, 0.3, 1.0, 2.0
To 2.0, 1.0 is closest.
To 1.0, 0.3, 0.2 and 0.1 are closer.
I wouldn’t call this evil, however: it’s merely demonstrating a technique that you should be aware of, if you’re a privacy-conscious person. It looks like they also provide some resources for avoiding stylometric detection.
I'd way rather have someone tell me "look at all the things I can find out about you" so that I can act accordingly (whatever that means!) rather than what we've mostly actually got, which is companies silently exploiting my data and doing everything they can to mumble reassuring but legally ineffective formulas assuring me that they deeply respect my privacy.
We should see it as an opportunity to learn how easy it is to associate different pseudonymous accounts. Nothing drives this point home better than a practical demo.
We can be pretty sure stylometry is used widely by bad actors already and we should not punish people who help to spread the word about these technical possibilities.
edit: maybe you'd catch some criminals if you tried to match reddit against dark web for example
Also think it's probably poor form to list users as examples without their permission.
Yes.
> This may be out of line but isn't pg on here with a different username, Levenschtein distance of one that's not included? Or is that just a very motivated 13yo account who writes a lot of admin-esque comments.
What other pg account are you referring to? I want to see it so I can see what my algorithm missed.
> Also think it's probably poor form to list users as examples without their permission.
You're right. I'll remove that - I just wanted some examples especially for people on phones who don't feel like typing. Thanks for the feedback.
HN offers many other threads which could be tied together, including:
- time of posting
- ratio of replies to top-level comments
- comments being mainly upvoted or downvoted
- sentiment (mostly angry, dismissive, questioning, etc.)
- most common topics (keyword analysis of post being replied to)
- ratio of new posting to post replies
- first-to-comment on a post
- lone comment on a post
- etc...
It seems very likely that sooner or later every pseudonym for posting content will get discovered and linked. The lesson here is don't post anything that would cause you undue shame or harm if linked directly to your legal name.
Native English speaker as well.
The second that I found out that requesting deletion of an account and its posts needed a MANUAL request to a single user (dang) I noped out so fast
But happy that the rest of you are still happy to contribute :)
Edit: From the "How to avoid .." page, there is the following sentence:
> Also, most authorship identification algorithms have poor accuracy when working with small amounts of words. This means the optimal strategy would be discarding an account either after every comment or after a small number of comments. Unfortunately, this is against HN rules and may result in a ban.
Can you clarify what this means and why it would result in a ban?
I have seen dang respond to users multiple times asking them to stop making new accounts especially but not always if it's to avoid rate limiting. I don't know if there's an official policy but it's definitely something I recall.
Imagine that for every new comment you want to post you would create a brand new account which you would use precisely once and never again. Then the stylometry would have just a few words and wouldn’t have enough corpus to get a reliable signature. If a lot of people does this it would be hard to figure out which account belongs with which human. ( Of course if you alone do this, your messages will stick out like a sore thumb. See xkcd 1105 )
> why it would result in a ban?
Because this practice is especially discouraged in the guidelines: “please don't create accounts routinely. HN is a community—users should have an identity that others can relate to.”
Maybe with some GDPR magic.
Drawing this into the logical conclusion, a user may opt to always post under a throwaway account, to avoid any possible tainting associated with a primary account.
Unless the author would run this against all HN user accounts, no need to flag the ones "of interest".
Very nice clean site, great work.
I guess it's difficult to evade it as the word frequency certainly catches all about the countries I frequently refer, programming languages, interests etc.
.. [0] https://hackaday.com/2022/10/20/render-yourself-invisible-to...
https://twitter.com/austingwalters/status/104189476543920128...
Made both Metacortex.me and insideropinion.com
The idea being you don’t actually need an active directory. It would drop in, figure out all the users (provided one account was on the AD) and would monitor everyone’s skill sets, morale, schedule, etc. Worked super well for what it was / is.
Out of curiosity: do you filter sentences than begin with ‘>’, indicating a block quote from another user? That might improve the accuracy a little here, if you don’t already.
I wouldn't rely on these results
You picked a user who posts a massive volume of repeat, template-y comments and found their former colleague who also posted piles of repeat, template-y comments, that being part of both of their jobs.
I picked dang as he is the figurehead of hn, and didn't want to inadvertently reveal some other user's identity.
At least the #1 close match (sctb) was a comoderator with dang, so they were kind of alts as the official voice of HN.
;-)
How long before large commercial indexers start offering an efficient (AI based ?) stylometry to agencies and states ?
wait... do you think the NSA is already doing this?
You just showed another possible use case for this kind of tools: "How unique is my writing style ?"
(Statistical stylometry is a little newer and more rigorous than manual stylometry, which essentially involved a human being's judgement call around the similarity of documents.)
Yields some results
This one seems pretty interesting
Does this ignore stop words? Or do all words have the same weighting? I wonder if only focusing on stop words would give a more accurate measure. Maybe we are more comfortable with certain stop words more than others?
https://en.wikipedia.org/wiki/Stop_words
"Stop words are the words in a stop list (or stoplist or negative dictionary) which are filtered out (i.e. stopped) before or after processing of natural language data (text) because they are insignificant."
2. There is a Russian mnemonic verse, which can’t be properly translated to English, at least it’s beyond my humble capabilities. It goes:
“Это я знаю и помню прекрасно:
Пи многие знаки мне лишни, напрасны”
The number of letters in the words give you the pi number: 3,1415… The meaning is: “I know and remember perfectly: too many signs (positions) of pi are useless and impractical”. Sometimes it’s nice to remember both things.
Edit to add;
Would be nice to have the https://news.ycombinator.com/user?id=username links included.
I am afraid to combine all these methods
Bit of a shame for useful posts/discussions.. but the internet is getting really.. finger print laden.
- Why is the author costco[0] not in this lookup?
- The text on that page is accurate it seems.
I ran it on Brian Armstrong's temp account from here, and it said it didn't write 10,000 characters:
https://news.ycombinator.com/item?id=3754664
EDIT: Or maybe it's something else because Brian only wrote less than 6k characters. But then why can my account be looked up?
Also, I would guess quoted replies are included, which muddies the analysis. Seems to be a very naive implementation. Much more can be done, but this was probably just a quick project.
A lot of discussion on the thread are over "how can we prevent this". I would like to know why should we not embrace this and similar technologies?
The benefits in my view are large - online behaviour tracks back to real life - and epidemiology speaking the value of millions of test subjects across every question are invaluable - from traditional medicine to "mass psychology recommendations"
I can guess some downsides (hiding from abusive exes) but am interested in studies, surveys, reports etc - any HN thoughts welcome
This is good to you?
Okay, let's just make it like China or SK where your login is your citizen ID and if you write bad things the bad word police will take you away.
Also, no, I have no alts.
But firstly the "governments will come and do bad things" argument - yes this is clearly and obviously a major problem - but not one solvable by technology in anyway. Fixing violent dictatorships is a IRL problem - one that requires enormous effort and sacrifices (see Ukraine for obvious example). We cannot pretend that a browser extension or a ground up rewrite of Twitter will defeat Putin or would have stopped Hitler.
As for "free" countries (something like 120+ have open free elections), we still have online abuse for voicing opinions that some people don't like (anything from pro/anti Trump to LGBT and bitcoin etc). Those are real consequences but rarely government inspired and honestly I suspect we need better support for police in prosecuting such things - I mean a death threat is a death threat.
In general my view seems to be we should have the same protections online as we do offline - and if those protections are "in theory only" that requires us to use our voting and other political power to chnage it - not to obfuscate IP addresses or so on.
The upside of tech is so great it is worth spending IRL to defend agains the downsides
>I suspect we need better support for police in prosecuting such things
We do see that! But mostly people on Facebook. Here we have had judgements of people who posted threats on Facebook because it is tied to your real name.
And yes, abuse is part of the "fun". Under your system, my 10 years old Leauge and CoD chats would have me locked up.
>I mean a death threat is a death threat.
Is it? I would find it more concerning if someone on the street tells me he is going to kill me than a kid on xbox live.
NOW there is a difference in systematic stalking and harassment online if I would get bombarded with DMs and messages to kys. I don't know how to solve. But a one-off comment is NOT equivalent. Then it feels like I'm just old? At 31? Is it really so serious?
My main point is not that we need to lock up everyone who makes a threat, but that we as a society will have to adjust our standards to the new normal.
Once upon a time every conversation was fleeting, every discussion in a pub or bar was ephemeral. Even Einstein and Dirac would walk home chatting without fear of being overhead. Then someone imagined it would be wonderful for the whole word to hear the erudite wisdom of those two geniuses of our age - and Facebook and Twitter and social media made it possible for every conversation in every bar to be captured and recorded and published - and we found out that Dirac and Einstein were just sledging each other and most other conversations globally were worse.
The new normal is that, like speeding, most evenings, conversations in most bars actually broke quite a lot of laws, from hate speech to sexual threats and basic politeness. And now the police can hear them as can everyone else - and discretion does not work on this scale - we either enforce the laws or change them.
That's a conversation for each judiciary- and likely to be either a balkanisation of the social media world, or a race to the top (we can all have twitter as long as we all behave to the standards of the highest / politest society. I am not sure where I stand on that.
Is it serious - hell yes. We are looking at a global technology with global benefits for all humankind - and if we want to communicate globally we need to agree what the standards for behaviour are on this virtual stage - from contract law to human rights and freedom of speech. We are inevitably going to build closer contacts - Brexit is a salutary lesson - and how we deal with freedom of speech online is just part of the jigsaw - but a telling part.
Until people can pierce the veil of your pseudonymity (which isn't all that hard depending on the platform and the person) and it isn't just online abuse and harassment anymore. "Tied to your real name" includes "tied to enough information about you that someone with plenty of free time can sift through various databases and piece it together" and most people have absolutely no idea how many such databases there are, and how much piecing someone can do.
I'll say something tangential: Even if we both agree that one-off assholes are largely inconsequential, and I think we do, such assholery has a broken window effect on a platform, where people see all the assholes running free and decide that it's either a place for them to be assholes or a place they should stay away from to avoid assholes.
I'm sure this is completely harmless and will not harm society.
The old posts issue is interesting- do you mean that there are posts from years ago you would find upsetting to be linked to you? Is this because you have chnaged your mind (a normal process society needs to understand) or because you said things thinking yiunweee anonymous that you would not have said under your real name? Far less of a social issue I think.
It does make for some interesting thoughts if we made everyone post under their real name.
The point that I am making is that it's incredibly easy to decipher why "track everyone under every identity they choose" can go wrong and lower the quality of discussion, and specifically, that it's so easy the fact you can't think of a single reason why it's a bad idea to completely eradicate privacy.
If I can find an alt of yours saying that you've quit smoking and then push tantalizing ads to you, you're going to bring me a better return than blind-firing into the American public.
If someone is looking for people who are easy to manipulate in borderline-illegal fashion (let's say, sex crimes), it's a cheat code if they see some throwaway account on HN comment on a post about the treatment of youth, "As a present high school student, I disagree with your statement because..." and track it back to a minor.
I also explicitly ask for real life examples and studies of harm - I can imagine and create examples but I much prefer the real world to my imagination as a guide. We learnt that as basis of science.
I also think there is a difference between privacy and secrecy. You seem to conflate the two - if your actions online were secret then advertisers would not send you smoking ads. Secrecy is probably impossible - privacy is merely the politeness of our neighbours. And at scale politeness is enforced - by social norms and sometimes legal measures. We are seeing this come in (GDPR) but it's hard to have legal enforcement before the social norms have arrived.
On the smoking ad front, Gabriel Weinbergs main argument is that searching for "red men's trainers" should be enough to serve ads without having to know if I am a 20 something graduate in wisconsin or a middle aged bloke in London. And I suspect he is right within a few percentage points.
As for online grooming -yeah this is a huge danger. Every parents nightmare. And still absolutely something that needs to be enforced in the real world. And may need extra police and social resources. But if we want to stop predators reaching out to vulnerable children then it requires co-ordination amoung many groups onleine and offline - funding, political will, training education over many years.
There will be no quick fixes for the problems tech is bringing - but I remain optimistic that the cost benefit ratio is worth it and that we can vote for and require change to defend against the dangers
which takes me back to my point - what are the real world examples of dangers so we can make sensible policy
https://stylometry.net/user?username=nickstinemates
Number 2 for me is someone I worked closely with for a few years, and then putting his name into this results in all of the people we worked with for a few years. So it seems content>style, or, we are all more alike than we thought.
Also, talk about a chilling effect. I was already vaguely aware of this, and now I'm overthinking every word I'm thinking/typing.
And yes, potentially very chilling. If you want to post truly anonymously, you might want to run your words through some kind of filter first.
I know their blog, which is their HN username, and this tool found their other account.
Perhaps ironically, this person stood out a lot because of this and I didn't forget them.
My guess is that as they commonly mention the project and I have on a number of occasions, that has formed the link. Plus maybe usage of common British terms, but that seems far less significant.
It's super interesting!
It would be good if there were more controls to filter the type of words and language that are used for the matching algorithm. So you could say exclude words not in the dictionary. I wander how that would effect my link with this other person.
Long live throwaways.
Is this? I thought that it was ok to have throwaway accounts, as long as they're not specifically to avoid a ban or something like that.
A question for the author (costco): You created that account in 2019 but you didn't post or submit a single thing until 4 hours ago. Why did you create an account almost 3 years ago for no purpose?
Does a low correlation with other users imply higher susceptibility to de-anonymization if I were using alts regularly?
Maybe we could be friends :)
Needs sentiment analysis IMO, otherwise you'll get "Here's a bunch of people who are JUST LIKE YOU", except they use a similar grammar style but hold opposite opinions on the same nouns.
Ok, fine, we'll present Idris with a fig leaf.
When I searched the other one, “soneca” was the first guess, with 0.4.
But when I searched “soneca”, the other one was not in the top 20.
Would translating to other language and back defend against this algorithm?
No but it would be easily adaptable especially given that Pushshift is archiving every Reddit comment. Based on some of the feedback I'm getting here I don't know if I should open source this even though it really wasn't that hard to make.
> Would translating to other language and back defend against this algorithm?
Yes. But then you have to send your original comment to a translation company so there are privacy concerns there too.
I'd say you should. I'd rather see this as being publicly and freely available to everyone rather than some shady "Big Tech" analytics company.
If the "weapons" exist, I would feel more comfortable knowing everyone can access them, not just an elite that can use it for their own (selfish) purposes.
Given the technique used, I don't see why something simple and local wouldn't defeat it? The "easiest" technique would be to use this weighting as a negative metric in rewriting.
There are modern offline translation systems available such as Bergamot https://browser.mt/
On the other hand, can we agree that this product is unethical?
In many cases, when a person uses an alt, it is a direct and strong signal that they do not wish their other posts to be associated.
So this product is circumventing the explicit will of the person, and making it available to anyone with zero effort i.e. there is no barrier to getting this info.
I met someone about 10 years ago who said they built this at a university. And their argument also was "actually this enhances privacy because it lets you know something something something". And yet their research grants were coming from one source only.
It can be used for good, but most often it won't.
It does create a high level of discomfort, because it illustrates well what privacy advocates try talking about to the population at large, but all that said.. how is it any different from regular scraping and analyzing it any other way?
This is a real question.
Imagine you get the urge to track someone, but in order to do that you have to spend a week writing some new software. That's a barrier. And because of it you may change your mind because it's a lot of work with little payoff.
But if that info is just one click away, it's a whole different ballgame.
No.
All those scared folks who naively think that it's not too late yet. Busted.
Running then the most similar person to my account did not put me in their top 20.
C'mon guys, work harder. That's not even close! :-D
Btw, I myself am only at 0.9999999999999999 so I guess I need to work harder at being myself.
Good job.
> Most likely candidates:
skymarshal: 0.9999999999999997
The other few usernames I tested (pg, dang, some random ones from this thread) all matched themselves at 1.0.
I’m a native English speaker as well, so I’m unsure how to feel about that.
looks like it can indeed
> Here are some frequent HN commenters: (EDIT: Removed due to privacy concerns)
How surprising that someone might object to being included in a demonstration of the erosion of privacy!
Is the site opt-in or opt-out?
* and I mean they author of the tool is here making posts, so I guess they have agreed to the TOS, but clearly someone who hasn’t agreed to it could also make this tool and scrape out publicly available posts without agreeing to anything.
I ever had only one account here and the closest match is at 0.47.
anyway, this app was able to identify a lot of my accounts. but a lot of the matches werent me. bold matches were almost all me. but i know there are many more matches than those that were listed. it mainly showed my most recent accounts.
i think most people would get a sick feeling in their stomach if they tried this app. i dont think people are prepared for a world where you can type someones name into an app like this and produce everything ever recorded online that was created by that person. not only this but everything highlighted and summarized to answer any question about that person. this is what advanced ai will bring us. an information implosion where the planet-sized ocean of data that is just floating all around us suddenly and violently coalesces into the objects of our new societal calculus. violent is a good word. and this is just the change that one can see coming with ai.
…Because I’m on mobile
Insightful that your personal experience and impact on you personally affects your decision. I invite you to think about the impact of the products you build in your CS career by putting yourself in the shoes of other people as well.
Some products should not be built, even though it's easy to build them.
And that accounts last several comments were flagged as dead.
I'm a native speaker, but my english succcccks.
Which user has lowest best match?
Mine is 0.58 so I'm really not that unique.
also maybe a tf-idf vector of top n words per user.
also could maybe do a same phrase analysis across the corpus to find some hand picked features.
timestamps could be interesting.
or, of course, let the machine do it with comment2vec.
If an account returns a high score for many accounts, does that also mean they’re relatively less original in style?
Is there a common open source library (Python, JS, whatever) that implements something like this?
Could they do much better?
I'd love to have the experience and or apparent wealth my "alts" have
One funny thing though, while your example says 1.0, for my own account it says 0.99lotsof9s4
Perhaps 6 or 7 digits is enough?
> uberduper: 0.9999999999999991
This is breaking anonymity that people incorrectly thought would not be revealed.
For some it might be awkward, others it might be quite problematic.
Analyzing stylistic similarity amongst authors
https://news.ycombinator.com/item?id=10050603
http://markallenthornton.com/blog/stylistic-similarity/
37 points by lingben on Aug 12, 2015
So, I don’t think it causes any new harm, if anything it gives you future risk aversion.
And I'm no lawyer, but it seems like there's also an outside chance of a breach of section 171 here as well, which is a criminal offence committed by a person who reidentifies de-identified data.
Plus - the laws have extraterritoriality. Vanishingly unlikely that you'd actually be pursued for it, but it's worth bearing in mind when you munge people's personal data.
The law seems to only apply where the deidentification has been made by the data controller, but HN admins changing someone's username, for example, if they ever do, would count. A person then using the tool to match another non-anonymous username to that account would seem to be caught.
Important to stress how much of a technicality this is, but that sort of thing can be interesting sometimes.
Holy shit, it works really, really good. It found all of my older accounts.
I can’t get over how phenomenal this is. Please put every one of your side project ideas into production!
One small nit from a user experience point of view..: it'd be easier on the eyes if you just truncated those cosine similarity scores (or whatever score you're using) after the, say, 5th digit. Showing the entire float is kinda messy to my eyes.
two spaces after a period
Whether someone uses an em-dash/single hyphen/double hyphens (which may correspond to house style they're used to)
Whether they use semi-colons
(Presumably harder) but consistent substitutions like loose for lose, break for brake, etc.
Use of accents
Fingerprinting certain linguistic traits and mapping that to time-zones as well as confirming there is a partial overlap in posts but never exact worked exceedingly well. Someone can't easily maintain a fluent conversation between themself on two accounts, but they can either get close, either through unnatural delays between sentences or just never interacting with the "other" party at the same time.
Woah
my current and my old account
This is to protect high profile users who are secretly enjoying programming in D rather than the language they are supposed to use.
And, of course, to protect users who feel they might be discriminated against if their background was known.
j_s,password4321,carolinew,colinwright,kuharich etc.
https://stylometry.net/user?username=j_s https://stylometry.net/user?username=carolinew https://stylometry.net/user?username=colinwright https://stylometry.net/user?username=password4321
Lowest match for j_s is 0.80 and all but one is black.
pg: 1.0
montrose: 0.604073065373204
mattmaroon: 0.5900372458160795
natsu: 0.5519832271289953
rauljara: 0.5418566694533273
waterlesscloud: 0.5378996309342633
damoncali: 0.5292014150349463
gruseom: 0.5290151637991445
kemiller2002: 0.5254174524920762
jfengel: 0.5231938496089998
jamesaguilar: 0.5229081613163672
houseabsolute: 0.5219738531025365
danssig: 0.5195368367601849
austenallred: 0.519343009683366
loewenskind: 0.5177030083877397
baguasquirrel: 0.5153841099708854
asdfasgasdgasdg: 0.5146704002447524
aptwebapps: 0.5144149629369845
allenbrunson: 0.512802806408646
danielweber: 0.5123620795710832You fail, I win.