AI unmasks anonymous chess players, posing privacy risks
science.org
science.org
https://axbom.com/keystroke-dynamics/
As early as 1860, experienced telegraph operators realized they could actually recognize each individual by everyone's unique tapping rhythm. To the trained ear, the soft tip-tap of every operator could be as recognizable as the spoken voice of a family member.> With straight keys, side-swipers, and, to an extent, bugs, each and every telegrapher has their own unique style or rhythm pattern when transmitting a message. An operator's style is known as their "fist".
> Since every fist is unique, other telegraphers can usually identify the individual telegrapher transmitting a particular message. This had a huge significance during the first and second World Wars, since the on-board telegrapher's "fist" could be used to track individual ships and submarines, and for traffic analysis.
Ham = amateur radio operator: https://en.wikipedia.org/wiki/Amateur_radio
This sounds incredible, to pick the right player out of 3000 candidates 86% of the time.
I am not sure that pruning the first 15 moves is enough to eliminate the information you get from choice of opening (which is presumably the intention of the restriction). For example, if a player religiously plays the Najdorf Sicilian as Black, you can immediately rule out many(most? ) positions that started with a French or a Ruy Lopez.
I'd like to see what the best results are from a model that just looks at the position after move 15, and use that as a baseline.
Reading the paper, they have an "opening baseline" which consists of frequency analysis on a player's first 5 moves. That model has 93% accuracy!
The mapping of first-five-move sequences to the positions obtained after them is almost a bijection (there are some transposition, but a small effect) so that's similar to my proposal.
I can't tell whether the 15-move cutoff is 15 half-moves, so 7-8 moves per player, or 15 moves each which is how every chess player would read the sentence.
Either way, I haven't completely read the paper yet, but I don't think it addresses the rebuttal of "I will just change my openings and the machine won't detect me".
Good point about baseline.
It takes 33 bits to pick out one human from everyone alive
Is it basically just 2^33 > ~8 billion humans, therefore that's the minimum information context to identify a single individual? But then what counts as an information bit - any valid Yes/No question? And how do you calculate the bit value of a piece of info (i.e. 3 bits for the knowledge of playing chess)?
The way to think about the information content of a problem or of something you learn is exactly what you're suggesting. If you numbered every living person on earth, it'd take more than 32 bits and not quite fill the 33rd bit.
If you then learn a person's gender, you can eliminate all the people with the incorrect gender, which is going to leave you either 31.x bits (assuming binary gender) or 25-27 bits of remaining entropy (assuming some non-binary gender and, say, a 1-3% incidence rate).
When the parent you're responding to says you get 3 bits for knowing someone plays chess, they're guessing that 1/(2^3) = 1/8 of people, in an undifferentiated sense, play chess. Of course if we knew someone's age or gender or country of origin, the conditional information value in knowing they play chess could be greater or lesser. And realistically no one is ever trying to identify a human among all humans (partially because it seems highly unlikely that there are many questions that could equally implicate the president of the United States and a six year old on the Marshall Islands in their answer). Each bit of information represents a halving of the entropy of the target surface.
I think you got to within 1 bit of the answer from first principles ;)
I'd assume that they are estimate that 1/8 of the human population plays chess, which feels like an over-estimate to me (but not absurdly, depending on your threshold of "plays"; by a similar process I'd estimate that at least 1/8 of humans alive are under 10 years old).
That's why I said at least 3 bits. If chess players are rarer, then knowing someone is a chess player is a stronger filter. (By the same token, knowing that they're not is a weaker one; but that's not the case we're discussing.)
I wonder how a chess GM would do at this test (although we’d have to restrict it to other top GMs that are active at the same time as them).
I don’t know chess but that sounds like a bad idea
As far as top players I think you need to be able to have variation in your repertoire if you play the same thing stubbornly you would risk becoming predicable and your opponent can prepare lines against you specifically. It might work for one tournament but afterwards people would have studied the games and developed counters.
Why? If you like it, you like it
For example, you can’t play any kind of Sicilian, whether Najdorf or otherwise, against 1. d4 (white’s second most popular first move, behind 1. e4).
When registering pick an elo. Provide a proof that you own the account in that Elo range and then you can create another account that will start in that Elo range.
So you get ONE pseudonym and then any other account you create are known aliases of that pseudonym.
And if you play in person, then any account/pseudonym you ever create are linked to your human identity as well.
Otherwise you can remain pseudonymous (not anonymous) for as long as you want.
But there is no way to do a mix of in-person and pseudonymous writing/chess/art/anything with a personal “style”.
The same thing with chess. Play chess under your real name (or write texts or draw illustrations whatever) and you'll never do so anonymously again. I'm not sure what's so upsetting or revolutionary here.
Show HN: Using stylometry to find HN users with alternate accounts
A (humorous) example: https://vc.blankenship.io
If you want to remain anonymous, use an AI filter for your written content.
If only there was some kind of re wording AI that can be run locally.
I'm happy to give it 6 months until the technology is there to do it, if it's not already—but, as long as there's an owner of the technology (that is, as long as the technology is pre- the point where I can easily roll my own), I'm skeptical of any owner in today's privacy climate intentionally forgoing the opportunity to suck up personal data whenever and however they can.
I thought "well this person seems a bit cynical" - you know, it's not a bad way to go outside yourself
yep, it was me I was arguing with, I'd written the previous comment a few years earlier from totally the other side of the argument :)
I don’t know why, but I don’t like that the headline frames it as a “privacy risk”. Are we really concerned about privacy when playing chess?
I think the world probably needs to accept there’s no such thing as “anonymous behavior”. Behavior itself is individualized. Therefore if behavior can be observed the probability that it is anonymous rapidly approaches zero with time and observations. The only way to be verifiably anonymous is to not be observed.
This means if you are a person at risk of harm if your identity is unmasked that you can’t rely on supposedly anonymous behavior. Bummer.
To the extent that headlines matter, I'd way rather that people worry about the privacy risks of de-anonymizing technology long before it's at the point where it's a practical issue. If we only worry about it when it becomes an actual issue already being, or about to be, applied to unambiguously privacy-invading matters, then, well, that's the way we've already done it and it's too late now—why didn't you bring it up earlier?
I'd also prefer to avoid the "what do you have to hide?" issue. Maybe someone, for whatever reason, does have something to hide; if they intentionally play chess anonymously, presumably they intend to do so. It shouldn't be up to me to decide whether or not their need, or even just desire, for privacy is legitimate.
(Of course, it's already too late—and has been since, at the very latest, the AOL incident—to worry about the onset of such de-anonymization, but it's always (or only almost always?) better to face inevitable future problems now, rather than waiting for that future.)
Who said anything about desire. I have no desire to unmask people.
To me it seems increasingly likely that observability and anonymity are mutually exclusive. I think it is hopeless because there's nothing you can do to stop someone from analyzing observations. We just never realized how much identifying information was contained in seemingly innocuous observations.
It seems like an obvious opportunity for a politician or even just a commentator, to pick up that baton. It's a well-paved path, everything is already written or you, and people really need some optimism and idealism right now.
Show HN: Using stylometry to find HN users with alternate accounts [1]
If someone would want to stay anonymous the unmasking % would probably be lower. The threat-model of the chess players doesn't include that they have to stay anonymous and need to switch up their way of playing.
I'd be very surprised if that actually works. Stuff like vocabulary can't exactly be turned off at will
I get downvoted whenever I say this, but anonymous speech is only allowed by recent technology and has never been a part of our ancestral environment.
1. The cost of printing/transcribing something 2. Literacy rates 3. Constraints tied to physical distribution
...the reach of that potentially "anonymous speech" was the tiniest fraction of what we experience today. And even then it wasn't necessarily anonymous, unless you just left books lying around?
Appreciate the link though. Thanks for engaging.
The reason you get down votes is its a relatively dangerous line of thought in itself that plays into the hands of authoritarians that would love to track every bit of information to an individual.
There may be some additional safety in conforming to the local "memespeak" dialect, and not using that dialect elsewhere.
In StarCraft II, having alts is tolerated (well, depending on your manners - nobody likes smurfing), but we have sc2revealed.com which takes crowd-sourced reports to try to unmask "barcode" (llllllllllll) players. Many pros try to practice anonymously on the ladder, because SC2 a game of imperfect information, and in a best-of-3 series (like in a tournament), you 1. don't want to use the same opening every match, and 2. don't want your opponent to immediately recognize what you're doing, or work on preparing a counter ahead of time.
How does this research help facebook?
Im very, very far away from Musk and his antics, but really some of those big companies seem to have lots of people who do passion projects.
Meanwhile an actual user has low if no chance to get decent support (probably for the cost of that of programmer they could get multiple people).
And yes I am aware that I sound anti illectual here and research the sake of reseaech can lead to nice things. I just think that the person will quit facebook to write poker bots and ruin the game for those who play it by detecting their weaknesses ( btw. I dont even play poker).
That's a risk when you pay workers to research or learn nearly anything new. If you run a pizzeria, you teach your workers how to make pizzas; they could turn around and make pizzas for another business instead. Maybe even open their own pizza shop and compete with you, using your own recipe.
I think accepting this kind of risk is simply table stakes for running a business.
For all we know, the guy who was kinda good at poker made the small break-thru that led Google's Alpha/Omega chess or Go achievement.
It's actually difficult to tell what piece of such complex systems are responsible for which - but in general, applying incremental piecemeal improvements have been monumental for the magnitude jump in progress in recent years.
For a time r/futurology was an interesting place for discussion, and there were really great comments to be found amidst the internet chaff .
One of the things I speculated about then was that ai doesn't need be sentient to ruin everything. Powerful tools in the hands of malicious actors could wreak havoc on the internet.
The internet could become compromised in so many ways via privacy invasions and data theft, aggressive spam, misinformation, propaganda and malicious code that nobody can reliably depend on it for much of anything any longer. One could receive a phone call from someone who sounds like their own mother, an ai that says things only a mother would know. That voice could be very persuasive.
that was the kind of stuff we talked about years ago. There were no instantaneous results, of course, and it became boring and uncool to keep going on about such things.
Nonetheless, it seems now that the tools and incentives needed to create a dystopia such as what I described are really starting to come into focus.
It's not the tech that's wrong, it's the populace in democratic states losing more and more power, to the point where most of these systems can only be described as hybrid regimes anymore. The actual power does not lie with the voters but corporations, the mega wealthy, their various lobby groups and corrupt politicians. There is no monopoly of power exercised by elected officials and law enforcement respecting the constitutions, it's different groups and fractions fighting each other for supremacy. AI is just another tool at their disposal, of course they're making use of it.
Guns can be used for protection as well as oppression. The internet can be about free information or about censorship and spying on users. We can live in digital Maoism or digital liberalism. It's up to the common people and for them to realize this before it's too late.
Other people want to be anonymous. Can they have beliefs and preferences independent of yours? Can their free choices have value independently of whether you agree? The value is in their free choice for it; your choice is equally valuable, but doesn't diminish theirs.
If you don't play like yourself you're likely cheating.
If someone suddenly plays a different style and also their move quality goes up significantly, that might be an additional indicator of cheating. All cheat detection works in a probabilistic fashion, since it is not allowed (and would be way worse) to actually observe players 24/7 in their home to verify whether they're cheating or not.
I don't see this changing anything, especially when we already know that chess.com is not impartial in its treatment of players.