How to break the 'rapper code'
blog.jgc.org
blog.jgc.org
The file basically uses a number of bytes between 128-254 to represent two-letter combinations. After an hour of tweaking a Python script I finally had the file decrypted. As someone who had never before dabbled in this sort of thing I felt very accomplished, although I quickly realized how rudimentary their methods were.
First I tried to find sequences of characters that, without modification, already looked like real words. I found one in particular that turned out to be the name of the city in which the game takes place ("Budayeen", though it was completely garbled).
That was a stroke of luck because that word appeared many times in the text and gave me some clue about the adjacent words, since there are only so many words you could reasonably put around the name of the city.
I tried globally replacing the characters in DISKTEXT.TXT but was ending up with a word half as long as I thought it should be (for "Budayeen"). One of the fortunate things that happened though was that it revealed to me, by accident, a couple other words (even though the decryption wasn't correct, it made some previously indistinguishable sequences look more like real words, which I pursued).
I think it really clicked that I was dealing with bigram substitution when I had to come up with a theory of how whitespace was so cleverly hidden.
Here's the lookup dictionary if anyone is interested: http://pastebin.com/pxJU7q7F
And here's the script that does the substitution: http://pastebin.com/8eYrGbzQ
You know, now that I actually think back to it, the lower 4 bits represent one character and the higher 4 bits represent the other. And there I was thinking in bytes the whole time.
I'm sure he means "relative to the typical criminal". The average criminal would probably not use a code at all, or would possibly use the classic A=1 B=2 cipher. Shuffling the numbers such that E=10, T=5, etc. is genius level compared to that.
You need to consider that this individual had to create the code and communicate its nature to compatriots on the outside. The code had to be simple enough for said compatriots to use in encoding information themselves, as well as being amenable to manual encryption and decryption. Since we can assume said compatriots probably were not exactly computer programmers or mathematicians. All of these requirements had to be met while keeping the code reasonably difficult for police to decipher.
I think it is, too often, tempting to only consider one side of the creation process in situations like this without giving due consideration to context. Giving full consideration to context, this would, to me, seem a relatively dangerous individual.
On another note, this is another example of how far law enforcement is willing to go, in terms of resources, to get their man or woman. Normally, the vast majority of us are not worth the effort of listening to our phone calls, or reading our emails, snail-mail, texts, or web posts. HOWEVER, once a spouse's body turns up, or ANY bodies turn up ... or maybe a bank goes under ... all of that changes. They will look through EVERYTHING. And you will be worth the effort to decrypt it.
It's not only national security that will get you that level of resource allocation.
If the police were in the habit of decoding simple ciphers like this, they could have put it in front of somebody that wouldn't describe this code breaking process as painstaking.
You can check out the modified code (with the translated solution as well) here: http://pastebin.com/umz5mM5F
The interesting bit is the underscore '_' --it can be either P or G depending on context, but I'm yet to figure out if there's some rhyme or reason to how it works.
Sure, some trial and error will probably yield results, but what he's describing is (part of the) systematic approach, not "overthinking".
You don't even need to know what a "trigram" is to solve single character substitution over English text; if you couldn't do it in a job interview in code on a whiteboard here, that'd be a "NO HIRE".
Also note that the "most efficient" ways of "breaking" (strange word to use when we're talking about Carmen Sandiego-grade ciphers) ciphers aren't always the best. For instance, the best way I know to break multi-character substitution in web apps is comically inefficient (in terms of ciphertexts required), but fits in just a couple lines of code.
When I did the New Scientist code breaking competition I just used an EMACS buffer and did M-% substitutions. http://blog.jgc.org/2011/06/how-to-break-new-scientist-ciphe... When I worked on the 'Reddit code' I think I just used EMACS again: http://blog.jgc.org/2010/12/breaking-reddit-code.html When I worked on how the Zodiac Killer enciphered the 408 message I started out by hand and then wrote a small program: http://blog.jgc.org/2011/06/how-zodiac-enciphered-zodiac-408...
For the hidden part of the GCHQ challenge that I discovered and reversed the key for I think I wrote some code in C: http://blog.jgc.org/2011/12/down-gchq-rabbit-hole-or-i-think...
In general, I like to stare at things, work by hand and write code to automate.
Or be fancy and use the solitare algorithm. It is pretty much designed for this case.
Though I might not still be alive in 25 years when he gets out of prison, I still didn't think it would be wise to automate the correction of his spelling and grammar. ;)
http://www.dailymail.co.uk/news/article-210 6384/Rapper-Kieron-Bryan-jailed-25-years-codebreaker-exposes-gangland-hitman.html?ito=feeds-newsxml
The @code array in jcg's perl script is a bit off. It seems his sister's name might be "Koh Koh" --probably pronounced like "Coco" of "Coco Channel" fame. I haven't found any supporting evidence, but it could explain the "6 25 4, 6 25 4," at the start of the message.