Zero-Width Characters: Invisibly fingerprinting text
zachaysan.com
zachaysan.com
Zero-Width characters tend to cause lots of issues when they are copied and pasted, which may alert a poorly equipped adversary that they're handling watermarked content. In addition, entities that are aware that you're watermarking text content in this way can just take screenshots of the text, transcribe it, or strip it of all non-ASCII characters.
The best solution I've seen is something I like to call "content transposition". The idea is that you take a paragraph of content and run it through a program that will reorder parts of content and inject/manipulate indefinite articles in order to create your watermark, while keeping the content grammatically correct. That way even if an adversary is fully aware that you're watermarking text content, they need two copies of the watermarked text in order to identify and strip your watermark.
The first barrier was that homoglyphs would inhibit text search, so I had to build an automated homogylph detection and substitution layer.
Once homoglyphs were stripped, the challenge was then to fit the entire search corpus into memory, so I compressed each page with LZMA, loaded it into memory, and decompressed on the fly when searching—probably not optimal, but still way faster than loading from disk.
I always wanted to try reverse engineering some of the watermarking systems so we could modify the watermarks on certain material, subtly leak it, and effectively frame adversaries while protecting our own spies in the process. Fortunately or unfortunately I never got around to that.
For forum posts, an adversary can do exactly what you say - just take the meaning of the forum post and write a cable on it.
Fleet pings are how massive fleets are formed in EVE. Basically they are a @here ping on Discord, or a server announcement on XMPP or IRC. Fleet pings are important because if you know when your enemy is pinging for fleets, you know when and where to find their fleets to fight them.
For that reason, the value in fleet pings is very time sensitive, and cannot be transcribed by a spy in real time. Usually orgs will set up a "ping relay" channel for its top fleet commanders, with a bot that relays pings from enemy alliances. This is a very common interception point for obtaining watermarked pings and burning spies.
Such an interesting game. I played for years and always want to go back.
One group decided that writing an automatic transposition program was too hard, and instead manually create transpositions for major posts. You can find the writeup here: http://failheap-challenge.com/showthread.php?16311-Taking-th...
Would love to see a swing taken at Pandemic Legion's new forum background watermark, it's a lot more interesting than the old one.
Some were identified and others... not
Months after my public comment, I got a phone call from an AT&T inventor who was prosecuting a patent on the same technology. They were very interested where I had read about the technique in the public literature. Alas, I could not remember where I had read that little factoid so I wasn't much use to them. It was disturbing to them that their patent claim was out in the open literature somewhere, but they could not find it.
Edit here we go https://en.m.wikipedia.org/wiki/Canary_trap
The Canary Trap, aka The Barium Meal.
My understanding was their leadership had a tool to help them select a few synonyms in anything they wrote, and a few synonyms was all it took to identify the account releasing the communication.
https://books.google.com/books?id=Vj-NDQAAQBAJ&pg=PA346&dq=c...
Red Sparrow also has a canary trap, and they refer to it thusly.
https://books.google.com/books?id=l-lnAwAAQBAJ&pg=PA428&dq=c...
See https://qz.com/1002927/computer-printers-have-been-quietly-e...
- Some characters affect word wrap in unexpected ways, depending on the script of the text.
- Some characters impact glyph rendering in minor ways. For example, the ligature between f and l may be interrupted.
- Some characters are outright stripped. For example, Twitter strips U+FEFF.
- Zero width characters often trip up language detection systems. I noticed that Twitter detected my English message as Italian with the presence of some zero width characters.
So it's not necessarily as useful as it seems. If you pick specific characters to strategically avoid these issues, it's hard to make the encoding very efficient.
Still, it probably has it's uses.
UTF-8 source code is nice for i18n, but it also opens the door to these kinds of attacks.
[1] https://github.com/NebulousLabs/glyphcheck/blob/f6483dd9e97a...
[2] http://www.unicode.org/Public/security/latest/confusables.tx...
Damian Conway is a wonderful mad genius.
Ideally, this would allow people to cryptographically sign comments and automatically verify comments that were signed by the same author, all without changing existing server software or adding ugly "-----BEGIN PGP SIGNED MESSAGE-----" banners.
One way would be to place N zws (zero-width spaces) between each original character, and treat each block as a digit which encodes a number. This could work for large original texts, but it would be clunky and very low bandwidth I think. E.g. if "." is zws, you could encode "fox" and the number 123 as "f.o..x...".
Better I think would be to create an alphabet with several of the zero-with characters and put the whole encoded number somewhere where it's unlikely to get trimmed or mess up line breaks (probably near the end in the interior but on a word boundary next to a space).
The hardest part would be making a transformation that wouldn't simplify it too much but would still be resilient enough to the transformations done by many forums like markdown/bbcode/trimming that the result could be perfectly converted into a PGP message. Maybe include some error correction?
Of course nobody ever said how we sign emails is beautiful.
I know this was done at one large tech company around 2010 for an internal announcement email. Different people got copies with slightly different Unicode whitespace, despite the email having ostensibly gone directly to an "all employees" alias.
The fingerprinting was noticed within half an hour or so. (Somebody pasted a sentence to an internal IRC for discussion, and as is the case with Unicode and IRC, it inevitably showed up as being garbled in exciting ways for some people).
However, the following countermeasure made be wonder:
> Manually retype excerpts to avoid invisible characters and homoglyphs.
Isn't this something you can automate? We should create linters for plain text (rather than code)? For example, depending on the language, reduce the text to a certain set of characters. Every character not in this whitelist is either replaced, or causes an error message the user (journalist) needs to deal with (i.e. remove it, or replace it with an innocent alternative, perhaps even proposing this replacement back to the linter project).
Of course, there are multiple ways for linting, which might become a fingerprint on it own. But then, if there are only 3 or 4 of such linter styles actually used (ideally, standardize on exactly one linting style), you can only tell which linter was used by the journalist, without any information about their source.
I was kinda on the fence, but I was considering my target: Journalists. A journalist may not notice simple differences like an extra space here or there when reading, but probably wouldn't retype a double space. I agree that this should be automated in some way, but it's a bit of an arms race.
There is a known variant when every subscriber of a confidential text receives a slightly different copy with the same meaning. But it's much harder to implement, and does not scale.
And this is without breaking into more intelligent heuristics where you swap out synonyms ("more intelligent" because you'd need to be careful not to alter passages that need to be kept verbatim, like quoted text or where a synonym might alter the context of the sentence.but with a little care I think that is achievable as well)
https://www.zachaysan.com/writing/2018-01-01-fingerprinting-...
Humans make fixes that are hard to codify in programming. I know it isn't perfect, but with my audience in mind (journalists) I thought it was probably safer than something automatic.
Regular expressions are powerful, and are used in plenty of production software.
If you can do the same manipulation in three ones of code it’s more likely to be correct and stay correct. And when you look at it again in six months you won’t have to stare at it. All those little time sucks add up as the code grows.
Edit: autocorrect got me twice.
Firsly, for relatively simple regex expressions, any competent developer should be able to grok them very quickly - at least as quick as the equivelant C#/Java/whatever code.
Secondly, it may be that regex is the most performant solution, and sometimes that matters quite a bit.
Honestly, I just don't get why some people are intimidated by regex.
1/ the regex engine
2/ host language
3/ problem you're trying to solve
But I've found generally I was better off not using regex for performance critical code.
HOWEVER (!!!) where regex consistently wins is development time. Not just writing the code, but testing (it's trivially easy to test regex) and updating the pattern matching (Vs updating the equipment character matching in an imperative language).
Yeah regex can get ugly quickly, but then so can any language if misused.
* RegEx is not very readable
* RegEx can be (very) slow
* It's not trivial to write RegEx code that achieves your goal in a high-quality way. Often quircks and edge-cases are missed.
I'm not saying that you should never use them, but oftentimes a (much) better alternative of achieving your goals is present.
See https://blog.codinghorror.com/regular-expressions-now-you-ha...
(I'm not sure what you mean by DoS attacks - are you referring to the exponential case of backtracking? If so, don't use a regex engine with that problem, and don't use lookbehind/lookahead assertions, which aren't needed to solve this problem.)
No, it is not a good solution to the problem; you ignore my earlier comments. English or latin text is not comprised of the sole ASCII characterset; it contains characters outside this set (quoting other languages, names, imported words for example).
Honestly, I do get your point about inappropriate use of regex, but this kind of simple text manipulation is well suited for regex. The biggest argument against using regex for this kind of problem is performance verses writing the same code programmatically in the host language (assuming you're using a fast AOT compiled language). However even that is a non-issue given the small quantities of text you're decoding.
Also I'd bet the regex in this instance would actually work out more readable because the transformations are basic so you're localising the text manipulation to simple rules rather than multiple lines of byte array reading and thus also potentially having to manually build in your own rudimentary unicode support too.
I suspect he put it low on the list because it could be a cat and mouse game of trying to anticipate all the potential information leaks.
It's also one of the things where a hex editor is extremely useful --- even if you're not working on low-level, seeing the bits directly can be a great confirmation of correctness.
Doesn't your editor handle that? I don't exactly remember but I kinda remember having a button to convert pasted characters to ascii (mostly used for those annoying stylized unicode quotes)
I've also found a lot of identical characters when handling Chinese text. Note that Google Translate does not handle these correctly.
https://github.com/pingtype/pingtype.github.io/blob/master/r...
It's about the Kangxi Radicals Unicode block, compared the CJK characters block. If you want me to write a blog post about it, please comment and I'll get around to it.
Even though some are noticeably different:
⿌ 黾
黾 is the simplified version of ⿌.
(In http://pingtype.github.io click Advanced > Regional, paste into the Simplified text box, then click "Simplified to Traditional")
So if you have cert and just want to copy paste the thunbprint to some file or application which needs to load it, then copying the full thumbprint probably won't work.
When I said fun I meant frustrating.
And don’t get me started on Microsoft and their fucking smart quotes...
Working with certificates on Windows in general is error-prone and difficult to automate (this coming from someone who spent more than a decade developing in .NET).
https://www.zachaysan.com/writing/2018-01-01-fingerprinting-...
One very interesting comment from an editor of The Weekly Standard.
It backfired when one of the top executives forwarded his own copy to the rest of the team.
https://www.cbsnews.com/news/should-management-spy-on-employ...
Copy and paste the examples into LibreOffice and you will immediately see them.
http://kb.mozillazine.org/Network.IDN.blacklist_chars
Or maybe black listing is not the best approach, maybe a mix of multiple approaches. First strip out stuff and then view the text in a program that displays "unconventional" characters? As a test I pasted the post's test sentences in Vim and the invisible characters are replaced by blocks of <XXXX> that are very hard not to notice. The more you think about it the more tricky corner cases you find :O
[1] https://gizmodo.com/382026/a-cellphones-missing-dot-kills-tw...
After all, we have many uses for the letter 'a', such as a) a*b=c and b) a as in apple. Should those 'a's have different Unicode code points?
Semantic meaning comes from context, and there is no context for a Unicode code point. Trying to insert semantic meaning is both a mistake and a patently impossible task. The article points out some of the wreckage attempting to do this causes.
Unicode does have separate code points for mathematical symbols. See for instance U+1D44E MATHEMATICAL ITALIC SMALL A. They see little use in practice though, apart from people using them for funky fonts on Reddit and such.
You mean like ⒜ or ⓐ or ᵃ? Not to be confused with any of ªa𝐚𝑎𝒂𝒶𝓪𝔞𝕒𝖆𝖺𝗮𝘢𝙖𝚊ₐ. Unicode has a whole lot of stuff that nobody uses…
There's also a separate Cyrillic а.
I guess we've always been in the days of "worse is better", like using CSV even though ASCII encodes characters specific to record separation[1].
[1]: https://en.wikipedia.org/wiki/Delimiter#ASCII_delimited_text
Granted, there are several failed experiments (e.g. interlinear annotations which are completely obsoleted by markup languages) and several pain points (e.g. shrug Emoji) inside Unicode. I don't really like them, you may not like them, but how would you make such a system without all these interim works?
Should the Greek alphabet be removed? Why have delta when d is just as good?
What I'm talking about is when the glyphs are identical, i.e. homoglyphs. Homoglyphs should be removed.
Tangentially related to this topic (ulterior fingerprinting), I wonder whether websites like Twitter might be encoding your IP address or account ID in pixels on your screen (so subtle that it's impossible to discern with the naked eye) to make it easy to track screenshots back to you.
A few years ago I had to fill in some translation keys within a ecommerce shop gui. Eventually, I had to revert some key to it's default value and therefore tried to copy and paste the displayed default value, but the system always refused to take that value. It complained that the string contained invalid characters, but I wondered because I just inserted the previous valid default value and to my eye there where not special characters anyway?!?
So I called one of the devs and after a few minutes he told me, that the output I copied contained a zero-width space and that character was not allowed by the validation engine. So when I typed the string myself everything went fine ;-)
Nowadays, I like to consult `hexdump -C` in such cases.
""" We're not the same text, even though we look the same.
We're not the same text, even though we look the same. """ example by cutting and pasting it:
https://jsfiddle.net/tim333/bjL018k1/
(It took me a lot of googling how to handle unicode in javascript, that)
Slightly to my surprise the zero width spaces survived being posted into this comment.
They were discovered a long time ago, and there are many other ways to hide data in documents.
Remember that you only need to encode enough bits for a relatively unique ID (and not unique for all files in the universe but only for files with the same content - for a low-distribution file, even 2-5 bits might be enough). On the application level, the most common applications and formats have a very large number of features you can utilize to encode data or simply insert it (e.g., Word, Excel, the PDF standard). On the bit level, unless the application vendor has invested in writing exceptionally tight, secure code, you probably can find someplace to hide/encode a few bits in a file.
But I think the author is on the right track with the solutions ...
> Use a tool that strips non-whitelisted characters from text before sharing it with others.
A more general solution is needed: Something that normalizes data in many formats, from text to Word to PDF to JPG to WAV to markup languages.
Personally, for non-security reasons, I'd love a utility that normalizes text to 7-bit ASCII (e.g., from UTF-8 characters higher than 7 bits) and that fits very efficiently into workflow (e.g., something that normalizes text in the clipboard if I press a hotkey combination). Anybody know of one?
Using the Option-Right Arrow to move through the text a word at a time has a problem though. The cursor appears to gets stuck at that point, and MS Word shows the font changing.
It might affect string handling where you split a string by space characters, and then compare words with a dictionary.
For example, Pingtype English which translates word-by-word to Chinese. (note the words "the" and "same" not getting translated in the examples).
https://pingtype.github.io/english.html
Google Translate handles it fine.
https://crashcourse.housegordon.org/coreutils-multibyte-supp...
"We're not the same text, even though we look the same.".split("").forEach(function(c, i){ if(c.charCodeAt(0) > 127) {console.log("Danger: weird character detected: code " + c.charCodeAt(0) + " at index " + i);} });
I'm not sure what to do with more complicated language mixing, but I suppose a script to tell you what languages are being used and where there might be "out-of-place" character codes would help.
https://marketplace.visualstudio.com/items?itemName=nhoizey....
- replace quotation marks with two apostrophes
- replace “I” with “l”
- replace parentheses with slashes or square brackets
- replace commas with semicolons
Each of these substitutions can communicate a single bit while being ASCII-safe and likely won’t change word wrapping opposed to synonyms.
Speaking of quotation marks, the non-ASCII quotes in your second item and apostrophe in the last sentence stand out too.
I dare you to inspect the html in Safari, Chrome, Firefox, and even curl, to see if you can view the raw character encoding.
Some you can, some you can't.