They Cracked This 250 Year-Old Code, And Found a Secret Society Inside
wired.com
wired.com
These unaccented Roman letters appeared with the frequency
you’d expect in a European language. But they don’t
represent letters—they mark the spaces between words.
It's implausible that these characters just happen to appear with a language-like frequency distribution and are all meaningless spaces. I suspect they actually have a meaning and provide a second message.To clarify, it's like taking "SthisEisCtheRfirstEmessageT" and assuming all the capitals just indicate spaces.
As a kid I was always making up codes and ciphers. Much more of a spy vs spy kid than a cops and robbers kid.
It's implausible that these characters just happen to
appear with a language-like frequency distribution and
are all meaningless spaces
Really? If I were to try to pick random letters I suspect I would end up mirroring the frequency that they appeared in English.Humans rather suck at picking random numbers. We skew towards picking numbers that seem "more random", whatever that means (which would probably work in your favour for picking random letters with an English frequency, though I'm not too sure), but we also avoid "randomly" generating streaks of numbers, because we feel those are "less random".
If you put two teams in a room, one flipping a fair coin (and writing down the results), and the other pretending to flip a coin but just faking the results, it is usually very trivial to pick out which team actually flipped the coin. They are going to have surprisingly long streaks of heads or tails.
I don't have any evidence to necessarily suggest it, but I suspect this anti-streak tendency will tend to be strong enough to interfere with any correct frequencies which may otherwise appear. ("Oh my, this is far too many e's in a row..")
Oh, heh, I can see it that way now. I had intended my comment to say that, since you'd be trying to reach that set of ratios to hide things, you'd probably fail miserably against any competent analysis.
Testing the "random distribution" like it was done - with a small sample size - is ineffective at best
* to put it in a less-polite way: how the F else would you solve a problem like this, with non-computational methods?
Well no, the linguist tried in vain to do frequency analysis by hand on ~88 symbols for ~100 pages for a couple months before saying "bugger this for a game of soldiers" and went on with her life.
"She tried a few times to catalog the symbols, in hopes of figuring out how often each one appeared. This kind of frequency analysis is one of the most basic techniques for deciphering a coded alphabet. But after 40 or 50 symbols, she’d lose track. After a few months, Schaefer put the cipher on a shelf."
This is an excellent article. When wired writes a good article, it is always amazing.
Another poster mentioned the Voynich manuscript. It's available on archive.org if anyone wants to try their hand:
http://archive.org/details/TheVoynichManuscript
Here's a list of others:
Take a walk down some of the older lanes in London, say near Borough Market or back up towards Southwark, or the other side between Brick Lane and Petticoat Lane, and imagine yourself back in the 1700s.
Coffee houses, close groups having meetings, private rooms upstairs in narrow houses. The feeling that true knowledge was being passed on. The meaning people found in the processes of the primitive technology.
It strikes me that the boring bits of the decoding (tokenising the symbols, entering the tokens) could be farmed out using a web site hosting scans of texts. The computational resource could perhaps be spare cycles on a PC with an appropriate application. Scope for lay science of a particularly interesting kind, and the refinement of algorithms as they are applied to a larger corpus of texts.
I feel kind of sorry for them, that at the end of their journey they found what was essentially a Rosetta Stone for the code they were decoding.
...maybe the symbold used as spaces are not actually random and there's another message hidden there, with another cypher, offering the writers of this "plausible deniability" regarding its existence: they could only give the way to decipher the first level of encryption and say that's all there is, while the really important information was hidden in the "space characters"...
(... now putting my tinfoil hat back in the closet :) )
You can analyse texts you believe to be similar (in language, period, subject, etc) to the coded message you are attempting to crack, and use that to build tables of these n-grams in various semantic units.
Of course, these are useful in many more things than code-breaking, and Google have various datasets they make publically available.
The Google books ngram viewer[2] is a fun tool to play around with, or for the more serious, you can download a corpus of ~24GB of analysed web data they've crawled (from around 1 trillion source words)[3]
One actual example of a code constructed in the manner described is the Playfair cipher[4] which was used for a time in the late 1800s, but is now thoroughly broken.
[1] https://en.wikipedia.org/wiki/N-gram
[2] http://books.google.com/ngrams
[3] http://googleresearch.blogspot.co.uk/2006/08/all-our-n-gram-...