Rot8000 – Rot13 for the Unicode generation
rot8000.com
rot8000.com
Λ̊1 → ⊻∪ά → Λ̊⋌
𝄞 → 뤔뷾 → 駴點
Edit: anyway, even with correct (a+b)%n it's plain bad idea.
Unicode is not English alphabet. Everything not in basic multilingual plane is broken automatically. And even in BMP there's going to be bag of glitches starting from hanging combining characters and ending to ‘oops someone normalised our string and it's now different’ (for site, not for user / Unicode).
한글: 0xd55c 0xae00 똼軠: 0xb63c 0x8ee0 霜激: 0x971c 0x6fc0 矼傠: 0x77fc 0x50a0 壜ㆀ: 0x58dc 0x3180 㦼በ: 0x39bc 0x1260 ㆀ: 0x1a9c 0x3180 㦼በ: (repeating)
So... yeah. Weirdness all around. Might have better luck doing this with some carefully crafted xor pad for each codepoint so that it's likely to hit a printable character but impossible to hit a character in the 0xD800..0xDFFF range (and similar ranges)... trying to "wrap" in unicode would require reinterpreting the codepoints to some continuous numeric representation.
[ArgumentException: Error serializing value 'ᄳᅳᅋᅁᅏტ㈣䳷ᅇᄹᄫ�' of type 'System.String.']
After realizing it was "?" that was breaking everything, I ended up with this round trip:"こんにちは。元気ですか。" → "ᄳᅳᅋᅁᅏტ㈣䳷ᅇᄹᄫტ" → "こんにちは。ጃですか。"
It's broken. I suspect Unicode requires more careful manipulation than OP anticipated. :-)
It also bypasses 32 control characters, technically making it rot7968, sometimes with an additional offset.
-> It also bypasses ⋍2 control characters, technically making it rot⋏⋬68, sometimes with an additional offset.I miss "obvious" stuff like that all the time.
Edit: silly formatting dropped my math punctuation.
AFAICS, it's actually using decimal 8000, not 2^16/2 = 0x8000, so I don't really understand how this is reversible at all unless they're just subtracting it back.
What we really need is rot88000h for the full U+0..U+10FFFF range. :)