Mimic – abusing Unicode to create tragedy
github.com
github.com
In Russia there is a government procurement portal. Where gov organizations have to post their requests to enforce competetion and best prices.
The usual tactics [1] of corrupt officials was replacing cyrillic (russian) letters with respective latin homoglyphs so only affiliated companies can find and win this contract.
[1] http://www.bbc.com/russian/rolling_news/2013/04/130409_rn_st...
:D
Maybe a low hit rate, but if you could automate it, you could run the scam on a lot of places.
Presumably this would trick users who would go check "C:\windows\system32\drivers\etc\" but not show hidden files. Seems like a niche subset, but still a neat trick.
As a concrete example, the following are fake links to Wikipedia (and entirely equivalent):
http://xn--wkd-8cdx9d7hbd.org (FAKE, same as below)
http://www.wіkіреdіа.org (FAKE, same as above)
It is true that network protocols encode these internationalized domain names in a subset of ASCII, but the user sees Unicode in his browser address bar or email. There is no restriction on how applications (like browsers) display domain names[2]; they can use Unicode if they want. This lead to all sorts of devious attacks[3].
[1] https://en.wikipedia.org/wiki/Domain_name#Internationalized_...
[2] https://en.wikipedia.org/wiki/Internationalized_domain_name#...
https://codepoints.net/search?q=https%3A%2F%2Fwww.google.com...
It's really difficult to do the right thing here. If Greek question marks share code point with semi-colon, it obstructs search and replace for question marks.
Subtle differences in how Japanese and Chinese are written has led to differently written characters sharing the same code point. It's nice that you can easily look up most Japanese characters in a Chinese dictionary and see how they are used in China, but it has become frustratingly hard to get subtleties in their written form right. The Chinese version may have the line strike through another line, while the Japanese only has it touching.
I honestly don't know how to go about posting how to same code points have different written forms!
But it seems like it would be nice if code editors warned about text outside ascii. You usually only want that in strings and comments.
Context is the key here. Greek text doesn't use the semicolon for other purposes and searching/replacing such single characters in source code is a terrible idea anyway (think comments, string literals...). So what is the prohibitive failure scenario here?
Indistinguishable (for humans) characters with different code points were a stupid idea, it's fine to abuse it in order to point out that fact.
A semicolon is not considered sentence final, whereas a question mark usually is. This makes it easier for software to auto capitalize. So at least that's possible a use case.
There's also the possibility that Greek requires the top dot to be square or circle, meaning it might in fact have subtle differences in print.
Such software will need and have a language setting anyway (e.g. for hyphenation). It doesn't have to and cannot rely on code points alone, so the characters (or rather, different uses of the semicolon) needn't have different codepoints.
And those people are a menace. "Ball bearings" is a single word with a space in the middle. The only way you're going to get reliable word segmentation is with a natural language parser and a lexicon with an entry for "ball bearing". At that point, you're already recognizing "aren't" as a word; the punctuation isn't really relevant.
On the other hand, if you're not particularly upset about messing up space-including words, there's no real reason to be upset about apostrophe-including ones either.
Case in point: 每 (chinese) vs 毎 (japanese).
I know that I am perhaps not well informed on this issue, but I also think that having multiple code points for what most people would perceive as the "same" character is not a good idea.
[1] https://en.wiktionary.org/wiki/%E6%AF%8F [2] https://en.wiktionary.org/wiki/%E6%AF%8E
Want to jump off ?
I'm sure a Japanese person would be perplexed as to why that extra "u" matters since it's the "same damn word".
入 vs 人
玉 vs 王
口 vs ロ <- my font might make those indisguishable out of context
千 vs 干 vs 于
of course in English we have 0 vs O and l vs 1
Compare 人 with: http://i.imgur.com/ZpFMQoU.png for one of the most obvious differences.
For Japanese there is: 篆書、隷書、楷書、行書、草書.
https://en.wikipedia.org/wiki/Seal_script
https://en.wikipedia.org/wiki/Clerical_script
https://en.wikipedia.org/wiki/Regular_script
https://en.wikipedia.org/wiki/Semi-cursive_script
https://en.wikipedia.org/wiki/Cursive_script_(East_Asia)
Can see more here:
There is a fairly good solution to this for things like code editors and URLs or search strings in browsers: If a string contains a non-ASCII code point that is a homograph for an ASCII character, swap the text and background colors.
This is true, but it's equally true that they both have undiverged continuity with 'Greek Α', which has its own code point.
It's a design goal of Unicode to support exact round-trip transformation with any code page ever in use in the real world, so they can't unify two characters that appear at different code points in a greek codepage without breaking things, even if they're always graphically indistinguishable in a font.
That's somewhat language dependent, though its true for most uses of most popular languages.
http://earthlingsoft.net/UnicodeChecker/
It offers a convenient utility to diff arbitrary strings, which is also quite handy for e.g. detecting normalization discrepancies, and installs a service so you can highlight a character in any app and use “Display character information” to see what it actually is.
I have Python command-line version in my PATH which displays the character info for arbitrary input strings: https://github.com/acdha/unix_tools/blob/master/bin/unicode-...
https://www.dropbox.com/s/9j9h5rjt4gu22hb/Screenshot%202015-...
Each hex value shown can be clicked to open the Unicode character info for that codepoint
For instance, \u2329, \uFE64, \uFF1C and \u3008 can be best-fitted automatically to \u003C (the regular '<' mark in HTML)
I work for a "major search engine" that does a lot of advertising & marketing stuff. To get the most out of it, we need customers to implement some javascript on their ecommerce sites.
As is often the case, javascript code that needs to get implemented on an ecommerce site often gets copy-pasted or emailed around a lot internally within a customer before it reaches the right person who can add it to the site's pages.
In this example somewhere along the way, a normal javascript snippet got all of the semi-colons changed from ; to ;.
In case you've not already spotted it, ; is not a ; but is actually "Greek Question Mark" (http://www.fileformat.info/info/unicode/char/037e/index.htm).
It was very confusing why Chrome was moaning about a semi-colon an illegal token. I had a genuine "Am I going mad? Seriously?" moment before I realised what was happening.
But yeah getting to that point relies on pure inspiration, in my case.
After fighting against word processor quotes, it's become second nature to double check it periodically.
:set statusline=%F%m%r%h%w\ [TYPE=%Y]\ [ASCII=\%03.3b]\ [HEX=\%02.2B]\ [POS=%04l,%04v][%p%%]\ [LEN=%L]
:set laststatus=2
Produces a status line when inserting and recording a macro like (the character under the cursor is 'm': ~/.vimrc [TYPE=VIM] [ASCII=109] [HEX=6D] [POS=0123,0020][67%] [LEN=182]
-- INSERT --recording
And, of course... `:help statusline` [ASCII=2>4] [HEX=0>4]
It's not ASCII, and its hex code is 6BCF.It also says "ASCII=252" when I put the cursor over "ü". Claiming that values over 127 are ASCII is just a malapropism.
For us emacsen you can do a (what-cursor-position &optional DETAIL) which is usually bound to Ctl-x =
I don't think it will be nearly this clean to add it to the modeline, but I'll take a look.
:syn match Error "[^ -~]"It protects you from these tricks by highlighting "troll" unicode characters in red.
After that it was a matter of just going through the usual steps of working out why something isn't working. I think in this instance I actually ended up manually retyping the code, and when getting "identical" code that worked that was when the penny dropped and I realised something was not what it seemed. If you paste suspect characters into something like http://unicode-table.com/en/ it will tell you right away what it really is.
Couldn't figure out why one of them wasn't working, and it was actually an audience member who figured out PowerPoint had turned a quote character into a "smart quote".
Why, if I may ask? If they introduce compile errors in your project, those should be caught by the CI build and test run, shouldn't they? In any case, accepting a code change just by looking at the diff and without even trying it out sounds like not the best course of action to me anyway.
"My simple piece of code looks perfect and should work without problems. Yet it won't compile! Help!"
Answer:
"Try running `./mimic --reverse` on your source."
Edit: Oh hey, I actually found one that fits the bill:
http://stackoverflow.com/questions/14925894/trouble-with-arg...
> ls | wc -l
and get > bash: wc: command not found
As it turns out, I need Alt+1 to type a pipe character in my keyboard. If I'm not quick enough releasing the Alt key, I'll type Alt+Space instead of just Space, which inserts a Non-breaking space[1] in Mac. This character is not a space, and therefore it gave me a weird "command not found" error.This lasted for months until I found out what the problem was - given that it was a combination of my keyboard settings and OS, finding the root of the error took quite some time. The hint? The "command not found" error had an extra space in front of the unknown command.
http://search.cpan.org/~sburke/Text-Unidecode-1.27/lib/Text/...
Mimic chooses replacement characters solely based on their visual similarity with ASCII. Unidecode, while still doing character-by-character replacements without deeper analysis, tries to optimize the replacement tables for transliteration of natural languages.
For example, mimic will replace Latin capital H with Greek capital eta (U+0397), because they look similar. However, Unidecode will replace U+0397 with Latin capital E, because Latin E is typically used in place of Greek eta when transliterating Greek text to Latin.
ps aux | grep foo
zsh: command not found: grep
It happens to me at least every other day.
I had the habit of pressing the ALT key a bit ahead of time before an OR operator.
if (foo || bar)
Only with variables, and almost always when using array indexes that are not 'i' or 'x'.
Sometimes is annoys me, and I've noticed that it's worse when I use CamelCasing and not as bad when i_do_this.
[1] https://tools.ietf.org/html/rfc5893 [2] http://unicode.org/reports/tr46/
Many programming languages support non-ASCII variable name characters now.
Just because you can do something doesn't mean you should.
It is usually worth keeping variable names and such in English in enable international collaboration. Also non-ASCII source files can get mangled in transit.
It happens all the time, everywhere, that people write code and stuff in their native language.
It happens all the time, everywhere.
08:11:32 >> Δv = 3
=> 3
08:11:39 >> p Δv
3
=> 3
My solution when I have problems like this is to start building a negative regexp in vim: /[^-a-zA-Z0-9 \[\]]
I then add other symbols as I find them. I can usually find the illegal characters in about 30 seconds this way—and I can add the non-ASCII glyphs that I expect to be present to my regexp.Substitute poison for healthy food.
Substitute poison with healthy food.
The first means you take away healthy food and give poison. The later means you take away poison and give healthy food.
[1] Except that "replace X for Y" sounds weird, except in the common phrase "replace like for like" (and probably some others!).
[1]https://www.chromium.org/developers/design-documents/idn-in-...
:edit+seems like the three umbrella unicode symbols are not supported on hn, are they supported in e-mails?
I'm not sure how frustrating this would be. Wouldn't most people just delete the character immediately and type a new one?
To my dismay, there were none. So I spent the rest of my afternoon correcting this glaring deficiency. Fellow Emacs users, protect yourself from Unicode trolls and grab it here: https://github.com/camsaul/emacs-unicode-troll-stopper
echo "hello world" | mimic --me-harder 100 | say“L-W-R-D”, “L-L-W-R-L”, “hell-erl”, “H-L-er-D””, “eor-D”, “H-L-L-erl”, “L-L-W-L-R-D”, “H-L-W-R-L”, “hell-W-R”
Compare with https://codepoints.net/ instead.
Edit: great, HN is broken too.
au BufWinEnter * let w:matchnonascii=matchadd('ErrorMsg', "[\x7f-\xff]", -1) > var ﷺ = 1;
< undefined