ROT8000
rot8000.com
rot8000.com
I call it "Mojibake Steganography": https://incoherency.co.uk/mojibake/
I think in principle (judging by the description of rot8000), my tool should be able to decode rot8000 messages natively, but it doesn't seem to work on the example given here. From looking directly at the codepoints given, I think the example is wrong. It starts:
u+7c5d u+7c71 u+7c6e - which works out to "]qn" instead of "The", unless I am misunderstanding something. And in fact that looks definitely wrong if we're expecting ASCII output because they're all more than 127 away from 0x8000, no matter how it works.
The rot8000 page says:
> It also bypasses 32 control characters, technically making it rotFFE0, sometimes with an additional offset.
I definitely don't understand how this is meant to work. Why does skipping 32 control characters turn it from rot8000 into rotFFE0? Should that say 7FE0? I still don't see how ASCII is coming out as 7Cxx.
Taking `char - 0x7c09` gets the expected ASCII output.
* control characters
* whitespace
* surrogate code units (U+D800 – U+DFFF)
1. https://github.com/rottytooth/rot8000/blob/main/Rottytooth.R...However the feature of rot13 (and rot8000) that you can use the same operation to "decrypt" it again is unfortunately missing in your variant.
One nice property of rot13 is it reverses itself; rot13(rot13(X)) = X. At least, for basic ASCII alphabet. Your UTF-8 encoding step makes that impossible. I wonder if there's a sensible Unicode-friendly algorithm that has that rot13 property.
But anyway ROT in itself is a pretty stupid idea anyway, usually just done for show.
> It is used to enclose the text in a sealed wrapper that the reader must choose to open - e.g. for posting things that might offend some readers, or spoilers.
AFAIK, this has been a common use of ROT13 since the 1980s. It also preserves substring search and message length (unlike BaseN encodings), which are occasionally useful properties.
In cryptography circles it seems to be kind of a running joke ("just use ROT13 encryption and you'll be set!" is something I've seen several times) ;) I know it was never intended to be secure.
But it makes sense then.
Of course if you work with ROT13 a lot, you will probably gain the ability to read it just by viewing the ROT'd code, defeating its purpose :) The structure of words also gives away a lot, since it doesn't affect spaces, capitalisation or punctuation. I still don't think it's very good at this usecase either.
How do you figure? It feels like the simplest way to handle eg spoilers in a universally portable and widely-recognizable way.
It actually skips whitespace, control characters and surrogate pairs [0].
[0] https://github.com/rottytooth/rot8000/blob/main/Rottytooth.R...
Running this on CJK text is an interesting exercise.
Base64 does lengthen the text by a third, which may or may not be a problem. On the other hand, it doesn't need special handling of control characters, and manages to hide word lengths well.
Bug was in a C++ base64 encoder component.
I bet using ROT on this will lead to unintended consequences because the original characters won't combine but the replaced ones will.
But anyway ROT is a dumb thing to do anyway so it doesn't have any real-world use.
ab => frowning face + brown texture = brown frowning face => ab
Almost all combining rules (including skin tone modifiers) require a zero-width joiner character between the person emoji and the modifier emoji. So really it's frowning face + ZWJ + brown texture = brown frowning face. (Although technically I don't think frowning face can be modified.) Also, there are relatively few ZWJ combinations.
Technically, there are some older combination emojis that predate ZWJ, mainly the flags, which are composed of two single-letter emojis, e.g. regional-indicator-U + regional-indicator-S = United States flag. So I guess it might be possible to get a couple of those.
And in any case, I think this page assumes that you're staying within the bounds of the basic multilingual plane (it mentions a self-inverting transform would be ROT32768), which doesn't include emojis or skin tone modifiers.
No, the skin tone modifiers apply directly to eligible person emojis; no ZWJ is involved. (Unless other modifiers that require ZWJ are also present, such as the gender signs.)
类籸籽籁簹簹簹簵 籸籷 籽籱籮 籸籽籱籮类 籱籪籷籭簵 籹类籮籼籮类籿籮籼 籵籮籼籼 籼籽类籾籬籽籾类籮 簱米籾籼籽 籼籹籪籬籮籼簲簷 籝籱籲籼 籶籪籴籮籼 籲籽 籿籲籼籾籪籵籵粂 籶籸类籮 籭籮籷籼籮簵 籪籷籭 籵籮籼籼 籽籪籷籽籪籵籲籼籲籷籰 籪籼 籪 籼籹籸籲籵籮类簶籽籮粁籽 籶籮籬籱籪籷籲籼籶簷 籋籾籽 粀籸类籴籲籷籰 粀籲籽籱 籷籸籷簶籵籪籽籲籷 籼籬类籲籹籽籼 籨籲籼籨 籪籹籹籮籪籵籲籷籰簷 籒 籽籱籲籷籴 籽籱籮类籮 籲籼 籪 籾籼籮 籬籪籼籮 籯籸类 籸籷籵粂 类籸籽籪籽籲籷籰 籬籱籪类籪籬籽籮类籼 籲籷 籽籱籮 籵籮籽籽籮类 籬籪籽籮籰籸类粂簷
Leet-speak for "plaintext".
E.g. stars and hearts get rotated but sunglasses do not
(EDIT: rewrote my example to use words because HN doesn't render emoji, duh)
> While rot13 is the self-inverse for a 26-character system, and rot47 for ANSI, the Basic Multilingual Plane of Unicode requires rot32768 (or 8000 in hex) for a reciprical cypher
Not all emoji is in the BMP, at least some are in the Supplementary Multilingual Plane.
It's weird to me that if you're gonna do this dumb "rot13 but for Unicode", you'd only do it for the BMP, and not ALL of Unicode.
Star = U+2B50 which is less than U+FFFF
Sunglasses = U+1F576 which is greater than U+FFFF
The details you might be missing is that some emoji existed in Unicode before color graphic "emoji" was actually a thing. The stars (and hearts) are examples of ones which used to be just a basic shape in the font but now are commonly full color graphical "images".