Can we believe our eyes? Misleading people with Unicode.
blogs.technet.com
blogs.technet.com
http://giorgiosironi.blogspot.com/2010/08/google-never-remov...
I used this to prank some people on the in-house SEO team at my last job. I'd ask them if they had done anything that might be considered black-hat. Then I sent him a link to a "site:" query on Google indicating that our site had been removed from the index.
e.g. http://www.google.com/search?sourceid=chrome&ie=UTF-8...
I'm thinking of a system that combines aspects of virus checking, malware detection, bayesian spam filtering, and spell checking.
A Unicode system can be supplied with tables of characters that could easily be mistaken (visually) for one another. These tables, combined with dictionaries, could spot words that could look like dictionary entries visually but are not spelled in the ordinary way. This approach could even spot things that have been problems for years in pure ASCII: confusion of 0 and O, of l and 1 and I, of rn and m, etc, It could also spot insertions of non-visible characters into such things as URLs and filenames.
Such a system would be able to spot .exe files that had names written in such a way that the .exe extension was not visually displayed at the end of the name. If you double-clicked such a file the first time, it could ask you if you realized that it is program you are about to run and not a ".jpg" as the name might suggest. In fact, it could ask you about any file whose real extension and apparent visual extension differed.
There will still be problems that will sneak through, just as today you can phish people with subtle misspellings that don't require anything more than ASCII.
But making this a part of the system's evolving general malware detection system, with human-created tables and heuristics borrowed from malware detectors, spam filters, and spell checkers, is the best solution, IMO.
How about the OS adopting the convention that any codes outside of a few trusted (expected) alphabets get displayed in a way that makes it obvious to a human that they aren't what they look like (eg, a bright red border or something).
How about the OS adopting the convention that any codes outside of a few trusted (expected) alphabets get displayed in a way that makes it obvious to a human that they aren't what they look like (eg, a bright red border or something).
AIUI, there are two major reasons this wasn't done in the first place, and why more complex solutions are necessary:
1. Those few trusted alphabets would probably include Greek, Cyrillic, and Latin, all of which have similar or identical characters with different Unicode code points.
2. The goal of Unicode support, localized domain names, etc. is for software to be equally easy to use for all languages, rather than to favor some languages over others.
That said, it might be advantageous to have a locale-specific approach, so that characters not used by the current language will be highlighted. But, that could be seen as hindering the ability of sites in one region to reach users in another region, doesn't work well for text that includes multiple languages, and malware writers will probably find a way to mark their characters as expected anyway.
Edit: also, the two words "get displayed" paper over a vast amount of complexity in the way operating systems and applications display text. It would probably be just as much work as any of the other solutions proposed.
Suggestions like this should get you a professional penalty, like a yellow card.
"Saw a specific problem, suggested an Are-you-sure dialog for this specific case, on top of that, one most people can't answer".
HOWEVER, what I find extremely worrying, is that the URL in the URL bar remained the same all the time, with the embedded query string site%3Anеws.ycombinator.com, and IT DIDN'T CHANGE TO ASCII when I changed the encoding of the web page. Haven't they/we learned yet, that this is a very serious security issue?
I discovered this as I was writing a paranoid HTML cleanup library and wanted to prevent the attack where a user sticks a text-direction-change character into the page and reverses the whole thing. As we've all just witnessed, that can't happen in a conforming browser.
But when viewing the page as a text stream, yup, it reverses and then never really unsticks. Everything's working as designed!
(Maybe my library should still restore the page flow after all... I never thought of how it could mess up view source. As attacks go, it's weak sauce... but like I said, it's meant to be really, really paranoid.)
Normally blacklisting is bad, but we're only targeting the few text-direction-changing characters that exist.
There are other such libraries for other languages, poke around. See for instance http://htmlpurifier.org/ .
edit: I get the same in Opera 11.50, Firefox 5.0 and some Chrome version.
Is Chrome and Safari broken or being responsible? (Are there settings to change the behaviors in any of these browsers...?)
So say your logical text is this:
ltr LTR.
where the capital letters are RTL chars (whether because they're actually RTL or because of an RLO in the char stream). That is, the above represents the reading order. Visually this would look like this: ltr RTL.
assuming that this paragraph directionality is left-to-right; a reader reading this text would read the letters in the order 'l', 't', 'r', ' ', n 'L', 'T', 'R' (which if you note matches the logical order, hence the name).Now say you mouse down between the 't' the 'r' and drag right until your mouse is between the 'R' and the 'T'. Those are your selection endpoints. But the selection happens on the _logical_ text, so what's selected is the 'r', the space, the 'L', and the 'T' (and even in that order). You can see that in the example above; if you mouse down between the two 'o' chars in "elgooG" and drag right to between the two 'o' chars in "Google", then you get exactly this sort of behavior.
Hi RLO elgooG
^ select ^
I expected (everything between the logical end points selected): Hi Google
^^^^^^^ (selected)
In other words, I expected text selection to obey RLO character as well.If that's not enough pain, read http://unicode.org/reports/tr9/ (the unicode BiDi algorithm)
If you're still liking it, I know a few people who'd love to hire you :)
Moving your mouse happens over the physical text and sets the selection endpoints. Then everything that's logically (as opposed to physically) between those endpoints is highlighted. Your "I expected" diagram is showing the text physically between the endpoints, not the text logically between them...
But I guess from now on I look at the chars to the left of the dot as well..
(This isn't a "look how cool command line junkies are" comment; I was just musing.)
This was ultimately stopped by Antispoof (https://secure.wikimedia.org/wikipedia/mediawiki/wiki/Extens...) but the bug reports are still interesting:
The difference between the two is that this phenomenon is not wholly in the past.
One of the biggest thorns of the situation is when editors bring up an official or semi-offical website related to the subject and using Я, pointing to its existence as "proof". No, that isn't proof; whoever is managing that area of the web properties is just a jackass.
Thankfully, Cyrillic as a poor approximation of the artwork doesn't always win out http://en.wikipedia.org/wiki/Talk:Superunknown#Track_name_.2...
And hey! It looks like discussion that I wasn't even aware of spawned on the Toys R Us talk page about it.
This was very common in Yahoo Chat Rooms when folks would pretend to be someone else by registering their name with the opposite of what they had (assuming it had an "i" or "l" in it).
They would then take a screen shot of their font and copy that exactly so they could appear to be the other person. I'll let you imagine the chaos that could occur because of this!
edit: on a related note, I half-jokingly tend to read RockMelt as rock-me-it…
(name changed to protect him, he was pretty peeved that I did this - though I think he registered a few places as pavellishir to exploit the r/n similarity, too :P)
http://web.archive.org/web/20080922234800/http://www.rnicros...
links
The attack described here is simpler: two unicode codepoints, roman 'o' and cyrillic 'o', usually look identical. So by substituting cyrillic we can make a file called 'hosts' that the operating system won't pay attention to. This is the same problem with punycode internationalized domain names, where paypal.com might be spelled with a cyrillic 'a' and mislead people. The fix for domain names was to restrict what unicode you could use where. I'm not sure what the fix is here, aside from always showing hidden files.
so you would say that there should be no file names using the cyrillic o? So if a russian-speaking person wants to save a file, that file name should be rejected? Or translated into a mish-mash between cyrillic and roman characters?
How will that work if that filename is reused on a system on which the default font doesn't contain the roman characters (I'm sure such a thing exists) and thus font substitution needs to happen?
The fix definitely isn't this easy. Maybe one could disallow homoglyphs of a different language than the one dominating the current file name. But this might be a lot of work and I doubt it's fool-proof.
> A utf8 decoder should reject portions of utf8
> streams that don't use the shortest possible encoding
so you would say that there should be no file
names using the cyrillic o?
I'm sorry, I was unclear. I should have said "don't use the shortest possible encoding for a code point". Cyrillic 'o' is code point U+043E while roman 'o' is code point U+006F. The canonicalization attack relies on overly liberal utf8 decoders that would allow multiple binary streams to be interpreted as, say, code point U+006F.This looks like the canonicalization attack, but is a different problem, one that is not solved by fixing decoders.
I would highlite the background of any character that is not from the users codepage in red.
for eg. if your local settings are us-en, any character not from that codepage will have a red background (or even in italics, some way to signify that the character is 'foreign')
Wide letters: A B C D E F G H I J K L M N O P Q R S T U V W X Y Z a b c d e f g h i j k l m n o p q r s t u v w x y z
Better looking letters: ϲ р с Ѕ І А В Е М Н О Р С а е о ѕ і ԛ
Anyone know of a better resource?
E.g.: "Архив.gz" (Archive.gz)
Perhaps limit system folders and files to ascii-only. Doesn't solve any of the picjpg.exe issues, but it's a start.
I doubt there is an acceptable non-heuristic solution.
1. Highlight any two adjacent characters from different languages. Virtually no one needs an English word with one Cyrillic letter substituted in.
2. Highlight any word whose letters all look like they are for the current locale but is made entirely of code points from another language.
But then what's the solution for someone who speaks both Russian and English?
CJK fonts often include not only CJK, but European characters as well. Almost always at least ASCII.
Ultimately, having a single font for all characters is desirable: Having to go track down more fonts because you're seeing � in your text is a pretty bad experience. Substituting other fonts is at best a kluge, as it often looks terrible.
Projects like DejaVu who plan to eventually cover all living scripts (http://dejavu-fonts.org/wiki/Plans) are not only a good thing but they are also making substantial progress.
Also, even in ASCII, there are a bunch of confusing characters (all depend on which font, of course): I (eye), l (ell), 1 (one), | (vertical bar); O (oh), and 0 (zero); {} (braces) and () (parentheses); 5 (five), S (ess), and $ (dollar), rn (r-n) and m; vv (v-v) and w; etc.
Wе nееԁ tօ fⅰnⅾ Ьеttеr ѕоⅼυtions tҺаɳ vіѕυаⅼⅼУ dіѕtіɳgυіѕҺing сҺаrаⅽtеrѕ.
The solution was to implement a regex of whitelisted characters; since it's an English-only program, this works well and is future-safe. For multiple languages, a blacklist is probably okay, but the difficulty lies in keeping the blacklist both complete and up to date.
It seems like it would be easy to implement a blacklist that automatically updates based on the current crop of registered names such that filters are applied to each registered name to generate a number of lookalike names which are all unavailable for use. This should allow any good-faith user to register any name they want without causing confusion.
I guess the solution would have to be in the terminal emulator? Would a blacklist of Unicode ranges be sufficient?
Change my system back to ASCII? :-)
I'm wondering if there's a way to make it more obvious that doesn't require running od. It should be immediately apparent any time there's a file with a whacky name, not something you find out two minutes into investigating a compromise.
I was kicked within 2 seconds from joining the server.
I have thought someone should compile a mapping of all the visually similar characters in Unicode to one canonical character. Then we can all include a check for potential impersonation in account-creation code.
Hosts file attacks are well-known enough that on windows I always set them to read-only, so that even administrators can't change them without first clearing the read-only flag.
facebооk.com is still available ;)
It seems like it would only work for hiding from people casually checking. Personally I'd open the file by typing the path myself, so I'd end up finding the trojan's file. The same would be true for any automated anti-spyware tool.
So yes, this looks like it would only affect a very limited number of people - technical enough to check the hosts file, but naive enough to do it manually and not notice the other hidden file.
Maybe because of the Control Panel revamp, or the fact that the menu-bar is hidden by default in Explorer windows?
I'm not sure which will look dumber as a consequence of this post: me or the Windows 7 UI.
attrib +h
at the cmd prompt, the file properties in explorer would have a checkbox for this too