Is Unicode Safe?
jefftk.com
jefftk.com
https://www.youtube.com/watch?v=ZPbyDSvGasw [DEFCON 21: DNS May Be Hazardous To Your Health]
Because I got so much traffic, I decided to make the sites useful, and I wrote a little php script that downloads the latest technology RSS feeds and displayed the headlines.
On one particular domain I was getting more than 20 hits per day.
The equivalent domain name issues are a lot tougher, and are going to require a character lookalike table or some other system of rules to warn the user.
your web browser literally downloads and parses documents all day long. pdf and .doc files might be marginally less secure, but they're by no means supposed to be executable.
if windows refuses to display the correct extension, I suppose you have to manually open them from whatever you want to view them with (File|Open) after downloading them to a certain folder.
I'd have zero qualms about doing this with a .txt file on windows for example - wouldn't you?
It'd not be any different on Linux if your email client saved attachments and marked them +x. Now as a matter of defaults, sure Windows is different.
Still someone could save file in different location and execute it. Is there a way to identify which program generated specific file? Or perhaps we could solve this by allowing these programs (like Skype.exe) save files only in specific folders. Is that possible?
Files should inherit the folder ACL, so a deny execute on a folder would work nicely, but end up breaking things for users.
Microsoft addresses these issues with the "WinRT" platform and sandboxed/appstore model.
One of the particular challenges with IDNs is that there are two versions of the specification, a deprecated 2003 version and the current 2008 version. For a few characters they provide subtly different transforms. The 2008 version also ratchets down on a lot of non-sensical characters — they are no longer eligible in domain names. The remaining permissible set is quite conservative to limit some of the issues seen in the original version.
Don't confuse Unicode with how to represent code points in bytes, please.
It's complicated. It's more complicated than any encoding standard that came before. It's also the most broadly useful, and the first standard to really take into account the complexities of human written language, as opposed to just one region's written language.
The right-to-left control character is for embedding e.g. Arabic or Hebrew script inside Latin text (or vice versa). It is actually a controversial feature of Unicode as some people feel it belongs in a higher-level protocol.
Check out the examples here http://scripts.sil.org/cms/scripts/page.php?site_id=nrsi&id=... for an idea of what rendering non-Western scripts can entail.
The potential for abuse is evident, but it seems like these primarily ought to be fixed in userland. For instance, by giving cues by highlighting characters in widely different areas (latin vs cyrillic) or by ignoring rtl for extensions when a string starts in ltr.
(Not to mention, in the latter case, if users are opening random docs attached to spammy emails, utf8 is the last of your problems.)
If you're not seeing a valid TLS session with a certificate signed by an issuer you trust not to allow these shenanigans, it really doesn't matter what chracters you're seeing in the URL bar.
Part of the problem is that a lot of the languages and tools we use pre-exist widescale use of Unicode and don't handle it very well. The Python 3 approach is by far the best one I've come across (would be interested to hear of other examples), and they needed to make a backwards incompatible change to handle it in a way that made it harder to screw up.
It is a complex technology, and inevitably there are going to be holes, but as in a lot of other cases, it is worth it (necessary, even), and as we move forward our tooling, languages, libraries and practices will get better and reduce the risk. The internet is a complex technology that can never be completely secure. Doesn't mean it's not worth it though.
It's like saying a dictionary contains dangerous information.
I think the problem is software that enables Unicode input but is not willing to handle all the different types of input. For example, it seems like a bad idea to even let people input combined words of different languages; that's why we have input methods that filter out bad combinations; and dumping this on the font renderer without making sure the difference is highlighted.
Countless languages borrow words from one another, especially names. For example, some people argue that calling Marie Skłodowska-Curie simply Curie misrepresents her character naïvely, as her name and country of birth were important to her. You could, of course, latinize it, but that gets you into trouble with the Turks and the Koreans and essentially all languages that had writing before the industrial age.
http://moc.lapyap.md2.shptech.com
(copy and paste link)
If it's just one RTL character then that should be fairly easy to filter out. Of course if that's a way a filter works then there will be other unicode characters you can add to the mix and still make it look the same for an average user and pass that particular filter.
One could identify unicode characters that belong to a particular character set (say latin) and see if some text contains more as one character set. Then invoke the filter if a text has more as 2 different character sets. Of course I can see that getting in the way of some use cases as well (text with translations in 3 languages for example)
hello backward world
http://philosecurity.org/2009/01/12/interview-with-an-adware...
The adware registered a key in the Windows Registry with Null unicode in the middle of the string so that the UI of Windows failed to display or modify that string.
Personally, I'd have preferred they disallowed it in the standard, but it's too late for that now. Anyone know why it was included (other than the obvious reason that it obeys the encoding rules)?
This can be tested by creating a dummy .txt file and changing it's extension -- the icon changes although file contents remain the same.
Is another unicode classic