Show HN: Send Secret Messages over Twitter as Public Tweets
github.com
github.com
if David Miranda gets stopped at Heathrow and:
On the following morning, the gener·al requested permission to return the emperor's visit, by waiting on him in his palace.
A pitched battle follow·ed.
But the pride of Iztapalapan, on which its lord had freely l·avished his care and his revenues, was its celebrated gardens.
...
is in his twitter account, it's in no way plausible that they're just innocuous tweets, and he can be compelled to reveal the secret.A true steganographic message would have looked indistiguishable from any other tweet that he would have made normally. this is a cute system, but it's not steganography.
So if Miranda doesn't usually tweet about the history of Mexico, he can pick other texts (written by him or others) which would sound more plausible as something he might normally tweet.
Having said that, the middle dot is not as unobtrusive as I would like, so perhaps it's better to rethink that part of the system, using some of the other suggestions in this thread.
you could imagine a system which uses entirely normal and habitual tweets, but encodes information in choices of synonyms used in the text, or whether or not punctuation was used in certain places, or the timing of the tweet's publication. lower bitrates, but plausibly deniable as to the message's existence.
The classic steganography methods you're describing may appear less obvious (for lack of a better term), but they're also easier to break once the pattern is discovered.
Then what you're writing will be functionally indistinguishable from random.
hmm ... but maybe too random, humans are bad at being truly random - the entropy in your period usage will be too high ... perhaps xor the RSA output with a OTP of something 'random' you scribbled yourself on a page ;-)
... seemingly random bits, xored with anything not directly related to the bits in question produces seemingly random bits... There are other ways of transforming a sequence of bits to look less uniform, though.
It's a text steganography app using a simple book cipher, written in Clojure.
I welcome any feedback from HN so let me know what you think!
Have you thought about trying to do a public-key based version? People could list their key in their profile.
I considered using public keys, but I didn't want the resulting encoded tweet to be illegible or look like gibberish.
The conceit here is that even a secret message looks completely innocent, something that no eavesdropper would notice as being out of the ordinary (though the middle dot marker does give it away, to people paying attention).
[1] http://www.imdb.com/title/tt0085334/
[2] http://en.wikipedia.org/wiki/Secret_decoder_ring#Messages
why not use a typo instead? switch the letter at that point to another letter (could be a different one for each sentence). use fuzzy matching (eg locality sensitive hash) to find the correct sentence.
and construct the dictionary from previous tweets rather than books. then tweets look like tweets.
Edit: oops, this idea was already mentioned by akkatirk elsewhere in this thread.
(This has nothing to do with steganography, but seems relevant nonetheless. )
https://www.cerias.purdue.edu/assets/pdf/bibtex_archive/PSI0...
The advantage of such an approach is that you can use coherent text/messages.
Additionally, @workmajj and I wrote TweetFS using Plainsight. It lets you recursively pack up directories and post them as an encoded linked list of Tweets to Twitter: https://github.com/rw/tweetfs
I presented Plainsight at Hack'n'Tell NYC in 2011 and a video was recorded: http://bit.ly/pecGgW
Plainsight uses each byte of the input message to generate tokens. Bits are used to decide how to traverse the token tree, weighted by frequency. The drawbacks are 1) verbosity and 2) incorrect grammar.
One of the lessons of writing Plainsight is that spam can be used to contain secret messages. Send enough gibberish to enough people, with your intended recipient included, and you'll look like a spammer--not a spy.
I also wrote a fuzzing tool, called Shag, to find edge cases, e.g. for single-byte inputs: https://github.com/rw/shag/blob/master/shag.rb
-- Example 1 (regular text)
Type your message to encode:
echo 'Meet at Union Square at noon. The password is FuriousGreen.' > cleartext
Then, pipe it through Plainsight: cat cleartext | plainsight -m encipher -f sherlock.txt > ciphertext
The output will be Doyle-esque gibberish: cat ciphertext | fold -s
which was the case, of a light. And, his hand. "BALLARAT." only applicant?"
decline be walking we do, the point of the little man in a strange, her
husband's hand, going said road, path but you do know what I have heard of you,
I found myself to get away from home and for the ventilator little cold night,
and I he had left my friend Sherlock of our visitor and he had an idea was not
to abuse step I of you, I knew what I was then the first signs it is the
daughter, at least a fellow-countryman. had come. as I have already explained,
the garden. what you can see a of importance. your hair. a picture upon of the
money which had brought a you have a little good deal in way: out to my wife
and hurry." made your hair. a charge me a series events, and excuse no sign his
note-book has come away and in my old Sherlock was already down to do with the
twisted
Now, decipher that ciphertext: cat ciphertext | plainsight -m decipher -f sherlock.txt > deciphered
cat deciphered
Meet at Union Square at noon. The password is FuriousGreen.
-- Example 2 (binary data) $ dd if=/dev/urandom of=/dev/stdout bs=1 count=10 | plainsight -m encipher -f 1984.txt
10+0 records in
10+0 records out
10 bytes (10 B) copied, 9e-05 s, 111 kB/s
Adding models:
Model: 1984.txt added in 0.89s (context == 2)
input is "<stdin>", output is "<stdout>"
enciphering: 100%|#####################################################################################################################################################################|474.67 B/s | Time: 0:00:00
which is a war is real, the proles used mind on the telescreen. He could see through all right to. You have read what said. 'Yes,' only in the MinistrySeriously, though, how would your library work for twitter?
It seems that the encoding process creates texts much, much larger than the original message.
One of the use cases of TweetFS is to use Twitter as a 'dead drop'. You'll generate a lot of tweets by doing that, but there's no harm done.
I don't think I hijacked your thread. The other comments also discuss textual steganography.
I was just kidding about the hijiacking comment; don't people know how to interpret emoticons anymore? :D
But this is getting a bit off-topic I suppose... :)
wget https://spamassassin.apache.org/publiccorpus/20030228_spam.tar.bz2
tar -jxvf 20030228_spam.tar.bz2
cat spam/0* > spam-corpus.txt
echo "The Magic Words are Squeamish Ossifrage" | plainsight -m encipher -f spam-corpus.txt > spam_ciphertext
$ cat spam_ciphertext
(8.11.6/8.11.6) 3 (Normal) Internet can send e-mails until to transfer 26 10 [127.0.0.1] also include address from the most logical, mail business for your Car have a many our portals ESMTP Thu, 29 1.0 this letter on internet, <a style=3D"color: 0px; text/plain; cellspacing=3D"0" how quoted-printable about receiving you would like width=3D"15%" width=3D"15%" border="0" width="511" Date: Tue, 27 Thu, 19 26 because zzzz@localhost.spamassassin.taint.org for
$ cat spam_ciphertext | plainsight -m decipher -f spam-corpus.txt
Adding models:
Model: spam-corpus.txt added in 2.57s (context == 2)
input is "<stdin>", output is "<stdout>"
deciphering: 100%|#####################################################################################################################################################################|543.84 B/s | Time: 0:00:00
The Magic Words are Squeamish OssifrageAt worst case after posting a bunch of gibberish Twitter bans your account.
If you randomly start posting "Four score and seven years ago" and the rest of the speech out of order you are going to get flagged.
I used to build tweet spam bots, I am very familiar with how they get detected.
I would have thought that the content of a tweet is quite low down on their detection list. IP address, user account, frequency of tweets etc would be much higher.
Instead, I was thinking that you might intersperse these in between your normal tweets, in coordination with the people you're communicating with.
How could or why would Twitter be able to identify these seemingly undetectable tweets as spam? They don't look like gibberish.
context: http://colorlessgreen.net/about
Now your message is encoded in the userids of who you retweet.
The receiving party would have to know the OTP to decrypt the message.
(Edit: This assumes you can get a list of favorited tweets ordered by when they were favorited. I just noticed the web interface orders them by time posted, not time favorited.)
If someone were to do this, it would effectively subsume all further "I implemented $X on Twitter" posts. Sort of like showing that a language is Turing complete.
There's no obvious way of telling where a "secret" sequence would begin and end.
For now, it might be best left as a coordination issue, similar to the choice of corpus: e.g., you know that you'll be tweeting secretly n times a day, at these specific times only, etc.