US Government plans to develop AI that can unmask anonymous writers
reclaimthenet.org
reclaimthenet.org
In the words of martha stewart, it's a good thing
I doubt TrueCrypt developers decided that it was Not Secure As Microsoft’s bitlocker after considering the paper that got released the following year.[0]
[0] - https://www.usenix.org/system/files/conference/usenixsecurit...
lot of fun ways to stick it to those bulk surveillance data lakes that may be the targets of such training sets
¹ There.
I write obviate then I re-read and replace it with “removes the need to do $foo.
Some writers at work forget that the goal of the written word is to communicate to the reader not impress them with word salad.
I regard it the same as clever code (note: not the same as complex code) - written for the author not the reader.
"Ernest Hemingway: he has no courage, has never crawled out on a limb. He has never been known to use a word that might cause the reader to check with a dictionary to see if it is properly used."
Ernest Hemingway:
“Poor Faulkner. Does he really think big emotions come from big words? He thinks I don’t know the ten-dollar words. I know them all right. But there are older and simpler and better words, and those are the ones I use."
No. Look, I had some "fight" in these pages two days ago because somebody complained of my style.
It is the farthest from legitimate to suppose that one's intention is to impress - that is accusing people of childishness, and you need to have solid grounds to suppose that. There are higher chances that one wants to be _precise_.
That would be precision in the expression of some content, to have such expression match as perfectly as possible a defined structure of thoughts you elaborated. You may want to describe the facts precisely, and/or the mental processes that those facts provoke.
Also some image subtly (and not so subtly) emerged, days ago, of "language as a cage": but if you want to express exactly some "/that/", it is the language that will have to provide the means - as opposed to a horrifying idea of "expressing what language allows".
There are contexts in which it is of paramount importance that the reader understands the message (e.g. instructions) - pay attention towards using the most digestible expression then. But when you want your content to be expressed as the defined sculpture of structures of ideas, while many different expressions are possible they are pretty much "what they are, because they are what they should be".
I would have sworn that was from Elements of Style but apparently it's a Mark Twain quote.
For me, the gold standard of this is Carl Sagan writings. Beautiful ideas, explained in the clearest language possible.
More precisely: not just having people less intellectually exercised¹, but also and especially creating some sort of "demand" that "things have to be simple". Non-recognition of complexity is also one of the cracks allowing populism. Your duty is to "spend one further thought", while some instances seem to defend an idea of an "as-if-constitutional right to reduced consideration".
(¹ I wrote the other day, «I also believe that "Now drop and give me twenty" will return us fitter personnel than "Please, be fed from a straw"»)
Most people do not have time, nor desire to do additional intellectual work after slaving away doing whatever it is they do for income.
Expecting most of the English people to know obviate is pushing it, for someone using English as a second language it’s just adding complexity.
It’s about aiming where you expect the average reader to be.
Because you know the word don’t fall into the trap of assuming everyone else does.
Obviate has several cognates in romance languages, so it is straightforward.
Now, the phrasal verbs that are so obvious to a native speaker, these are the real head-scratchers.
Or street slang. It may sound like dumbing it down to you, but it actually adds complexity for a non-native.
"Register": I immediately know what it means and it takes tenths of a second for me to make the decision to click (or not to).
"Sign up": despite having a CEFR C2 level of English and using it at work every day, I still need to think for a few seconds if this is a registration or a login (cf. "sign in").
For a native speaker I suppose the latter is (marginally?) easier, but for a non-native, it's much harder, not even close.
“Create an account”
“Subscribe to mailing list”
“Submit contest entry”
It is one of the downsides of being a native speaker. I write pretty much like how I speak. It's hard to use it like a foreign language, but that's what technical international English feels like to me, in a way. In my experience, fluent second language learners master that style more easily than native speakers do.
”Comprehension” is not some audience-independent property of a text. Comprehension is what happens when a text is well-calibrated for an audience. If your audience is unlikely to know the word ”obviate” (which seems true of many engineering settings, where the audience is international), you absolutely improve comprehension by replacing it with words the audience is more likely to know.
I agree with you that ”obviate” sounds better, and in a setting like a blog post, monograph, or even an hn comment, that’s probably what I’d use. But in e.g. a work email for colleagues in another country, you reduce the risk of miscommunication by following OP’s suggestion.
There lies the spring coil. The less it is established what your audience is, the more you shift from "calibrated expression" to a form of expression which is as absolute and universal as possible. (See my previous post here, at https://news.ycombinator.com/item?id=33045225 .) If you have a message for an audience, you will try and condition its expression according the audience; the more the audience is abstract, the more the expression will be unconditioned, as if intended for an ideal audience.
That said, there is of course a gray zone here.
Occurs to me there's another point in your favor as expressed in Orwell's version of a well-known passage from the Old Testament 'Ecclesiastes' into modern English. He makes the point brilliantly. Aesthetics does come into it!
https://www.orwellfoundation.com/the-orwell-foundation/orwel...
That isn't detecting who wrote it, but certainly is the first step.
For some value of "identified". However good it is, it can't definitively nail someone as the author of a piece of text.
Or compose + (-) (-) (-) for (—) which is wider depending on the font.
I wouldn't doubt for a second that NSA is already using shit like this for their global passive adversary thing
If this works it's pretty much the equivalent of a mandatory state ID on every online interaction. If it doesn't work very well, then it's going to be that, plus the risk of randomly being flagged.
As a society I don't think we have anything to gain from it; it's certainly tech that's put of Pandora's box, but that doesn't mean we shouldn't use all our cultural/legal means to prohibit/control it.
Anyone in their basement could build this today, but it only becomes a problem if that person has control over police and intelligence forces. Which means the potential for good regulation is bigger than with other tech.
Edited to add point about regulation
Maybe that's the point.
It could also be announced because they plan to use it publicly soon, to "prove" that some person they want to get for political reasons is the same as some evil terrorist/pedophile/serial killer.
This is 100% going to happen, so my guess is it can’t be authoritative. It’s likely to be used to whittle down a list so that humans can review the results.
Feels like the beginning of a hybrid Minority Report + Enemy of the State movie. I’d watch that. I don’t want to live in that world, though.
> This is 100% going to happen, so my guess is it can’t be authoritative.
Why do you think that it won’t be authoritative?
That of course has its own risks, if you assume full government access to providers and unlimited surveillance capabilities. If you are fully paranoid you could even run your text through a local instance of (less accurate) Apertium[0] for the first few rounds, or even for the whole process if you find the result is different enough from your original.
An offensive one would be to post fake offers to trade government/military intelligence on the dark web writen in the style of the politicians backing up this measure, so that they are put on the feds list and thoroughly investigated.
Seems like an impossible task. There's just not enough signal to noise ratio to make that distinction. Not if the goal is to distinguish among thousands of people.
Among tens or hundreds, might work if they don't take any countermeasures.
"sir you are under arrest for maybe thinking about banning dairy production, which is at conflict with national security. We found an anonymous text online that has the same idea."
Why even bother to come up with a pretext then?
It would be trivial to mask text and change irrelevant opinions to be anonymous again. Though the real problem is related to posing as others.
Reminds me of ilillliillilill. There's a reason people use names like this.
AFAIK there is quite a bit of examples from security labs where malware authors aren't necessarily identified but at least fingerprinted based on naming conventions, patterns they use across multiple projects etc...
That sort of fingerprinting could expand to correlating someone's anonymous software projects to other examples of code elsewhere (ex: if they contribute to source available stuff).
re: the example project you mention specifically, it does feel like using tools like that almost as a linter for natural language would be a fingerprint in itself.
EDIT: As far as OPSEC goes, a fun tidbit. A friend of mine identified a PR I submitted anonymously to them, simply because of the style of PR comments I made.
A paper got released in 2015 that claimed 94% accuracy in identifying authors.
I’m sure it would have been quite easy for the NSA to figure out who Satoshi Nakamoto was too considering PRISM was also scooping up everyone’s email.
"It is a far, far better thing that I do, than I have ever done; it is a far, far better rest that I go to than I have ever known.",
or this from Pride and Prejudice:
"However little known the feelings or views of such a man may be on his first entering a neighbourhood, this truth is so well fixed in the minds of the surrounding families, that he is considered as the rightful property of some one or other of their daughters."
He expertly massages the worn keys to type out a lengthy and witty response about how he'd imitate the style of a famous novelist whose prose exists mostly to result in the sacrifice of as many innocent trees as possible - demonstrating how to stretch "I'd use Dan Brown's writing style" into nearly two full paragraphs of text.
cf [1]
"I think what enabled the first word to tip me off that I was about to spend a number of hours in the company of one of the worst prose stylists in the history of literature was this. Putting curriculum vitae details into complex modifiers on proper names or definite descriptions is what you do in journalistic stories about deaths; you just don't do it in describing an event in a narrative."
[1] http://itre.cis.upenn.edu/~myl/languagelog/archives/000844.h...
lel
If anyone is looking for an impactful project, Tails OS maintainers and it's author academics seem receptive to bringing it onto that privacy-minded platform:
https://gitlab.tails.boum.org/tails/tails/-/issues/5726
The New York Times found Mr. Lyons by looking for writers who fit those two criteria, and then by comparing the writing of “Fake Steve” to a blog Mr. Lyons writes in his own name, called Floating Point
Sent chills down my spine.
I also pass it into Hemingway[0] first to make my text lean and non-superfluous.
https://news.ycombinator.com/item?id=33034918
Am I the only one seeing the irony and contradiction here? Not the same people at all -- subgroups at best , but the Director of National Intelligence is part of the administration. Perhaps I am missing something -- feel free to comment -- I am curious what everyone thinks.
I remember hearing DARPA was actively seeking research in the field around that time. In principle, I'm not absolutely against my software being part of the chain of events that leads to the decision to kill someone, but I don't trust the US government (or, realistically, anybody else) to independently verify an identification made by such a system.
I'd be surprised if the three letter agencies aren't using something at least as good as what I wrote by now.
I cross checked using statistically improbable words, which helped confirm or exclude weak matches.
We're flexible on that format. Anything legible and relatively easy to translate into. Call it ANONSPEAK.
It will probably be, aesthetically, horrible.
Every verb is done in passive voice, punctuation is added - wherever possible - to make sentences appear more complex than they need be, and of course there is an effervescent use of sesquipedalian terms where shorter similes would otherwise suffice.
Trouble is if you are a revolutionary leader of some kind you are probably going to be saying new things that no-one else talks about - which renders both anonspeak and the ai detection kind of redundant.
I guess the application for this then is in the interim to stop people or online groups becoming revolutionary by tracking and deradicalising them with targeted manipulation.
You don't even need a 'writing fingerprint', you just need to parse comments which reveal identifying information such as 'I participated in project x', 'I taught at university y', 'I invested in startup z'... then when you combine all the identifying information, you can narrow down the pool of possible matches to a single person.
You could probably do it with just basic text matching, no AI required.
Now they'll copy China
I find it very funny
USA next decade:
"Posting bad things on twitter reduces your credit score"
Don't confuse that with "social credit" systems, whereby China prevents you from riding trains if you say something naughty.
https://mronline.org/2022/07/27/national-security-search-eng...
that's why the US wants to ban TikTok asap, because they don't want china to be able to do what they are doing for decades too
> it would harm them because Twitter posts are unlikely to represent a meaningful variable when predicting someone's creditworthiness.
people get fired and arrested already for posting stuff on twitter, in both the US and Europe, so no, it's not just just a "twitter moderation" thing
Ask yourself why they are allowed to exist and still operate despite unable to grow and are loosing money for years, talk about anti-competitive practices, unless it's in reality a government body in disguise
If you believe that, then you also support efforts to force big tech to respect freedom of political speech, yes?
It has begun right here in the United States and we just are oblivious to these new dastardly form of social credits.
That's my understanding as well.
> it's a convenient talking point in the west
When online comments get worked up about The Social Credit System, a key thing I believe they're trying to do is spread awareness of how disturbing it is that a government is even considering such a thing that, as we understand it, is closely related to being a core technology in an authoritarian dystopia.
While it's not implemented at scale, the unnerving fact is that govt policy makers did a careful enough take on a social credit system to decide that it was worthwhile investing (probably non-trivial amounts of ) money and resources into exploring it and did eventually reach a point where they were a handful of steps short of wide-scale implementation.
Try and keep up.
Where exactly does it say it’s ever progressed beyond trials and announcements ie implemented at scale - oh it doesn’t
https://nhglobalpartners.com/china-social-credit-system-expl...
>As of December 2020, more than 80 percent of all the provinces, autonomous regions, and municipal cities had issued or were preparing to issue local credit laws and regulations.
> Now they'll copy China
Rinse and repeat.
Failing to develop that technology, leaving it to China or some other authoritarian state to do, would be more likely to harm liberties, wouldn't it?
Seems like you're just "damned if you do, damned if you don't"-ing, no offense.
Why are you making this up?
Such as, you write a comment or an essay, feed it in and it just dumps all your styles, idioms etc and makes all your stuff sound bland and normal.
Seems like it would only work if you had a targeted population and a writing / style sample ffor a lot of people.
People have been doing this for decades.
The key point here is that it's AI-driven and at scale -- in other words, mass surveillance.
Personally, I see this as part of the US IC's mission, despite the potential domestic detriment.
(it's rôle and &c., goddamnit)
Of course, the article just says they're planning to do this. It doesn't say it's going to work, and our closest examples, forensic science like handwriting and blood spatters and polygraphs and all that, generally don't actually work.
Then you get to accuse anyone of anything.
Tho I guess I could just use gpt3 and say hey can you write this for me in your own words or summarize. Or a few rounds of google translate roulette
https://www.jstor.org/stable/30204514#:~:text=They%20were%20....
This is 2017 https://towardsdatascience.com/hamilton-a-text-analysis-of-t...
I believe we need countermeasures, so we need to create algoriths that add fuzzyness to our writing to protect privacy.
Perhaps it is time for Creole
Stylometric analysis did suggest a single person on that list. The easier thing for governments to do at the time would have been to just spin up a node in the first year and look at the IP addresses.
He had no desire to become known back then and likely never will. It's only more dangerous now compared to the threat before of being locked up like the LibertyCoin guy (who just got released a year ago).
NS is happy to stay in the shadows, nearly everyone respects that decision especially in a world of crypto scams and ponzis. Surprised they never linked the domain name purchase to him though.
It's known as your "hand".
In an era when privacy has become hugely diminished under the hands of both governments and corporate interests it raises the question of what rights to anonymity anyone has in either a public or private forum, and at present there's little if any consensus on this which ought to signal that any such project is premature.
Unlike yours truly—who usually speaks his mind irrespective of whether he's known to his audience or does so anonymously—many will not speak their minds out of fear of being ridiculed, or humiliated, or exposed, or out of the risk of offending—risking the breakup of a friendship, etc. Same goes for whistleblowers whose public utterances, if not done anonymously, usually costs them their jobs.
If people fear that their autonomy to act in an anonymous manner has been removed then they're unlikely to act at all, silence being the better part of discretion.
This would have huge negative repercussions for society, our institutions and our governance—after all, the secret ballot is one of the cornerstones of our democracies. If we're not careful AI could undermine the ballot by unmasking what users think or how they actually vote and it's not hard to see how this would lead to coercion thence totalitarian government.
That said, in this world of widespread almost instant communications, actors who intentionally act out of bad faith can do widespread damage, especially so when they do so anonymously. Knowing who they are would minimize the damage they are able to cause.
Similarly, in a distantly-related post on HN a few days ago I referred to the increasing loss of respect for our important institutions and for the way we're being governed and how I thought that faith could be restored. There, I suggested that as a part of that process we need to unmask the hidden processes of government and that this would also include the naming of those who originate policy, law, etc.:
"If we're to restore any faith in our governance then this protection [hiding originators of policy] must stop. Decisions made by government employees must be open to public scrutiny, similarly, the origins of government policy—laws, regulations etc.—must be traceable back to its source (those who initiated said policies).
Systems without accountability will always become corrupt."
Thus, there's a real dichotomy at work here. For some things anonymity is essential, at other times it's a curse. And from the many recent instances of where the gnomes within government haven't acted in our best interests then I'm damned sure that putting AI to work here won't bode well for us either.
I've little doubt that the technology will be abused, and by virtue of the fact it will automatically silence a large proportion of the population who need speak out and who should do so anonymously in the interests of all. Even if they aren't targeted directly just knowing that there are systems in place that have the potential to expose them would be sufficient to silence many—as AI analysis of their words could be used to determine their identity at any future time (living with ongoing stress from potential exposure of one's ID would likely be unbearable for some).
Given past history and current bad behavior of governments in these areas, I do not believe that it is possible to put such a system in place that would gain the full confidence of all players involved. It would have to have sufficient protections locked in place to provide full public accountability as well as having inbuilt mechanisms that would ensure the system could not be abused by governments. At present, such conditions cannot be realistically met—not by a long shot.
Before anyone or any entity could let AI loose on this project and simultaneously state with all honesty that sufficient protections were in place for the project to proceed with safety would require many other prerequisite protections and 'safety measures'—which currently do not exist—to be incorporated (locked) into our governance. For instance, a whole raft definitions and concomitant laws pertaining to privacy are needed—and that's just for starters.
No doubt this project will proceed without those prerequisite protections, ipso facto, it will also be abused.
PS: note my quoted point about government policy etc. being open to public scrutiny. Here such questions arise such as where did this idea originate, what are the names of its instigators and what are their motives for instigating this development—not to mention others such as what are their qualifications, experience, etc. (perhaps, given the enormous potential of this AI application to damage society, we may even need to pose questions concerning their political beliefs and allegiances).
It's no accident that this information is missing with this announcement.
Voynich manuscript ---> US Govt AI stylometry machine ---> 42Furthermore I consider games like poker or Magic the Gathering unplayable, that is the extent to which there is literally absolutely no privacy.
But don't mind me, just got lobotomized is all.