With unicode, I really don't know what to do. What strings need to be sanitized (or validated)? File names? Urls? How do I sanitize them without causing agony for most of the world? Are there other unicode attack vectors?
With unicode, I really don't know what to do. What strings need to be sanitized (or validated)? File names? Urls? How do I sanitize them without causing agony for most of the world? Are there other unicode attack vectors?
The general answer is to not allow (e.g., ignore it, remove it, but do not render it) in any identifiers, filenames, domainname labels, usernames, and other security-sensitive strings, U+202E between characters of scripts that are not bi-directional and which flow in the same direction. More generally, do not allow mixing of scripts in certain contexts -- for example, do not allow mixed scripts in a single domainname label. Some of these rules need to be implemented by, e.g., DNS registries, and may need to be tailored to their specific needs (e.g., you might find that .kr has to allow mixing of ASCII letters and Hangul characters because it's common in South Korea to add "ing" to names to make brands out of them).
UTR #36 doesn't go far enough, in my opinion.
[0] https://unicode.org/reports/tr36/
[1] https://www.unicode.org/reports/tr36/#Bidirectional_Text_SpoofingThe Trojan Source paper has a decent treatment of the area too, seen in a different application: https://trojansource.codes/trojan-source.pdf.
In the article it is shown how this can be misused to show executable files as a different file type:
file[U+202e]fdp.exe shown as fileexe.pdf
~~~~~~~~^^^^^^^
| |--- read this right-to-left, i.e. exe.pdf
|----------- this is the Unicode character Right-To-Left OverrideIf you "sanitize" Unicode you're just removing that functionality. Remove that character, and the right to left reversal wont' be there, right?
If you have to break it in the name of security and sanity, it shouldn't exist in the first place.
I am continually baffled about why the HN comment sections are such a hotbed of anti-unicode sentiment. Perhaps most of y'all don't have to write software that deals with non-English text in any significant way.
If you just didn't have buffers you wouldn't have to avoid buffer overflow attacks! Come on.
Yes, this aspect of SQL is poorly conceived because the only way is to manipulate it with text processing.
Parameters should be inserted into a query language using AST manipulation, like a Lisp backquote.
Now within a particular program, we can wrap something like that around SQL generation, and solve the escaping problems in one place.
That approach doesn't apply to character data, which has to be literal and primitive in many situations.
Some things are meant to be simple, and Unicode is massively complex. True, these are rare-ish corner cases, but every programmer knows it’s only ever a matter of time until you hit any possible corner case.
I’m not saying, screw people whose languages don’t fit into ASCII (I’m one of them btw), but I really don’t want to know the vagaries of Unicode representation to inspect a file name.
Maybe Unicode needs a “safe mode” where only outright characters are allowed? Or maybe the unavoidable complexity is a design flaw, for some uses anyway.
I feel like we’re close to someone discovering Unicode is Turing-complete...
Operating system should already know if a file is executable and desktop should be using this information when presenting files instead of conventional names.
Yes it would mean that when you download a file you have to set the executable attribute before you can run it. It's already almost exactly like that on Linux. That's a good thing.
That already happens. If you set the "Details" view in Windows, it will display one file per row with various columns showing file metadata, including the file type. So for the example in the post, you'd see "[Word icon] annexe.doc - [the date] - Application" as opposed to "[Word icon] annexe.doc - [the date] - Microsoft Word 97 - 2003 Document".
In linux it's pretty obvious: https://www.rootusers.com/wp-content/uploads/2018/01/gnome-t...
The main advantage is that the UI for opening documents (vim file) is different from the UI for running executables (./file) which is different from the UI for running executables as administrator (sudo ./file).
This makes it impossible to accidentally execute a file that looks like a document.
If I download a pdf, I don't want it to run. I probably want to open it, but if I've gone through the trouble of downloading it onto the filesystem as opposed to just opening it in the web browser, there may be other things I want to do with it. Maybe I want to print it, or maybe I want to break it up into component parts. Or maybe I want to transfer it to a sandboxed environment before I open it because I don't trust it.
If you don't want to take an active role in maintaining security on your computer, then by all means use a commercial operating system. Pay a license and let Microsoft and Apple make those decisions for you. They have competent people who can program reasonable default settings for the average user. If you're running Linux or an open source operating system, you've already decided to take these responsibilities into your own hands.
That is a beyond ridiculous take. The last Windows OS that did not make me mad was Win98, and every time I have to use Mac machines I want to gouge my eyeballs out. I want an OS that lets me shoot myself in the foot as much as I want.
> If you're running Linux or an open source operating system, you've already decided to take these responsibilities into your own hands.
no, I just want to use the fastest possible OS for my use case on my hardware, which is Linux.
inotifywait -m -e create "$(xdg-user-dir DOWNLOAD)"|while read d e f; do chmod +x "$d/$f"; done
You can add "-r" if you want to also do it on subdirectories, and you can specify another directory instead of the default downloads directory.
Though I don't really recommend doing this since you should not be executing downloaded binaries often enough for this to be worth doing.
I literally do this multiple times per day. What in hell do you folk do with your computers?
Thanks for the script ! There's a few tens of thousands of files in my downloads folder, I hope this isn't going to pollute the inotify fds too much...
Hopefully not leaking company information to clever attackers.
I work with a number of clients in high security environments where you have to get permissions to perform operations like execution (or even setting chmod x on a file). It really does limit people running random things and causing destruction.
I have no idea in what situation I'd need to load so many executables via the browser.
Executables tend to come compressed with a lot of other files nowadays, and compression software knows to set the executable bit when extracting them.
cd ~/Téléchargements; chmod +x <try to remember the first letter><hit tab 15 times><miss it><go check inside the firefox download manager><restart><end up typing the whole name>
before being able to run it.The Windows UI could just render that character as a tofu box, and not evoke its reversal semantics.
Consider writing a bash script that makes a subdirectory with current date if it doesn't exist, downloads url from the command line argument into that directory, changes the permissions and runs it.
Then your workflow could be:
- keep terminal open
- Alt+tab to firefox, browse, copy link to executable file
- Alt+tab to terminal, ./handle.sh <paste the link> ENTER
- Alt+tab to firefox again
- rinse and repeat
There's many benefits - you don't get thousands of files in one directory (which can lead to performance degradation and even random errors), there's little risk that you run the wrong file, you don't waste as much time.
If it's from a GUI, file browsers have the capacity to that that bit too.
> Maybe Unicode needs a “safe mode” where only outright characters are allowed?
What is safe for me may not be safe for you. I need RLM characters in my terminal emulator [1], you probably don't. And who knows, maybe RLM characters could prove to be an attack vector, by e.g. displaying filenames or other text backwards.Every Unicode character and code point exists because somebody, somewhere, needs it.
Not quite. That and because some twits are willing to cave in and add it, because adding crap to Unicode gives them a sense of purpose which makes them blind to the harm they are causing.
Every Unicode character and code point exists because somebody, somewhere, thought that somebody would find it useful.
Sure, just like they benefit from knowing the difference between 'V' and volts; 'J' and Joules; 'K' and Kelvins; 'A' and Amperes; π the letter, π the circle constant, and π the exotic particle; 'm', meters, and mass; 's', seconds, and position; 'g', grams, and gravity...
假設所有的意義都有自己的符號,那就太好啊
The real issue are the thousands of characters that appeared incidentally in some ancient text, either as typo or as a weird interpretation of a common character, which then ended up in the Kangxi dictionary, and then subsequently imported en-masse into Unicode.
Example, 𠒇 - the only known use (in non-ancient times) of this character was being the official name of 雷莊𠒇, winner of the 2017 Miss Hong Kong Pageant. According to her, she intended to write 雷莊兒 when she applied for her official documents, but somehow the officials interpreted it as 𠒇, which is really an archaic form of 兒 (at best). Reportedly she's changing her name back to 雷莊兒
Thousands of such characters exist, if you look at the page where 𠒇 is supposed to originate, more than half of this is obsolete -- https://www.kangxizidian.com/v1/?page=125#gv
(So yeah, you're correct in essence but picking on 囍 as an example probably doesn't really get your point across...)
囍 is used frequently, but it is not part of any writing system and does not convey any linguistic message. The opposite is true for 𠒇 - it is not used frequently, but it is part of a writing system and is used to convey linguistic messages.
To presume a character is "real" merely because it exists in the Kangxi dictionary is as valid reasoning as presuming a character is "real" because it exists in "some other encoding standard" that you've been dismissive about. It's just that Kangxi is the de-facto Han character encoding scheme before computer encodings were invented. (Unihan even contains all details about the radical, stroke and even page number where it was sourced from)
A lot of those characters appeared once in some ancient text, and took the Kangxi form due to transcriptions from scribes across the centuries, but we actually have no evidence that they are "real" (at any point in time). Some of these characters are known alternative forms of common characters, or are only known to appear in some ancient text before Han characters were standardized. Some are plain typographical errors. It's like a 3 year old child learning to write "ABC", which looks a bit weird, and then the unicode committee assigned 3 code points to them.
Those are entirely valid for Unicode. Han unification in Unicode is already considered a mistake. That's why "unified" code points now also have explicit, higher-numbered 'equivalent' code points that unambiguously refer to a particular graphical form. The graphical form is the whole point of Unicode.
It's used by being hung on the wall, like a painting[1]. Like I said, it has no linguistic use, and thus it is not part of a writing system, which puts it outside the stated scope of Unicode. It is the exact equivalent, for weddings, of the upside-down 福 character that is hung for New Year's, or the wreath that Americans hang for Christmas.
But it does not correspond to anything in any language; there is no Chinese sentence whose spelling would include 囍. Note that the 新华 dictionary entry says "Double 喜. Generally used at happy occasions such as weddings.", and there are zero examples of the character being used. ( https://zidian.aies.cn/NDQ4MA==.htm )
The wording of the 新华 entry is almost identical with the beginning of the 汉语大词典 entry, and it's instructive to quote the rest of the entry:
> Character used at happy occasions. Commonly called "double 喜"[2]. Generally used at occasions such as weddings. Often cut from red paper (or gold leaf), or written on red paper, [then] pasted onto a door, window, or wall, in order to indicate a happy occasion.
[1] Actually, a much closer comparison would be the traditional magical talismans that use elements from the writing system in a freeform way to express various desired goals. https://en.wikipedia.org/wiki/Fulu
[2] This dictionary is nice enough to make the fact completely explicit that 双喜 is the name of the character rather than a definition. 汉语大词典 doesn't even attempt to provide a definition for this character.
Literally all emojis are not part of a writing system nor have proper linguistic use. Who on the earth write emoji by hand or even spoke it? Unicode already broaden the scope for so long, whether it is good or bad, nobody cares now. And beside that, a lot of symbols are also culture dependent. If you are not used to live at there, you just have no idea what it actually means.
Yes, that's true, and they've been a constant source of problems for Unicode ever since the decision was made to let them in. They stand in violation of Unicode's declared principles and purpose.
Note that 囍 is not considered an emoji by the Unicode standards, though that's what it is in fact.
> Who on the earth write emoji by hand or even spoke it?
But there is an example of this - the name of https://en.wikipedia.org/wiki/I_Heart_Huckabees [wikipedia title uses the word "heart", but the actual title uses the symbol] was frequently spoken aloud.
I had to look them up so I might as well save others the time.
It's definitely an extreme view[0], and I'm open to being argued out of it, but I really think we should use a different and much simpler encoding for the vast majority of things even if it doesn't perfectly reproduce the written text of the language. Writing systems encoded the language in a way fit for the medium of paper, and maybe we need to consider computers a different medium and adapt our language to them differently too.
[0] and probably anglo-centric.
The problem is honestly just that unicode is used to convey things like "this is a PDF" or "this is going to execute code" and it's attacker controlled. It's a terrible UX that's been abused for as long as paths have existed.
So I guess hope that operating systems will find a better UX for "this executes" and that they'll stop "a process executed" from being a game over situation for users.
They just turn your url into punycode, and more recently display that instead of the raw url: https://en.m.wikipedia.org/wiki/Punycode
Technically, punicode is the raw URL and old enough browsers will display all non-ASCII domains as punycode. It was only for a short while that browsers naively decoded punicode without restrictions.
Wink-wink ;)
Besides refusing to render U+202E when it is between characters of scripts that aren't bi-directional and flow in the same direction, the UI could display the string in ways that make it clear (e.g., different color, maybe add a warning tooltip, maybe add a dialog around an operation that could be dangerous, or maybe refuse to perform dangerous operations).
Most Unicode attack vectors seem to target people, they are ways to obscure from a person the true intention of what something is. As a developer it’s our job to try and mitigate this through how we display the data, warn about risks and block particular attacks. So it’s about working out how to prevent a user from making a mistake.
It seems to me in this case obviously Unicode became “a thing” after the development of a file extension. Clearly if it had been the other way around the file extension would have no meaning, maybe that’s what we need to move towards, store that metadata elsewhere on the file system.
Fundamentally Unicode from and “untrusted” origin shouldn’t be trusted in an executable context.
One could even translate the file extension into something more user-friendly (e.g. .exe -> Application), making it a "file type" column. Which is of course what Windows Explorer has been doing for decades already. Of course, that results in everyone turning the classic file extension back on...
The other problem with hiding the extension is that you then have "file name" collisions, two files will appear to have the same name but only differentiate by the "type". I think that's wrong.
Eg. use my libu8ident. For technical details see the C++ proposal to use secure identifiers: https://rurban.github.io/libu8ident/doc/P2528R1.html
> With unicode, I really don't know what to do. What strings need to be sanitized (or validated)?
Just yesterday I filed a bug on the KDE terminal emulator [1] for stripping away too many characters. I'm sure that the fine devs had no idea which characters are safe to leave in and which are not. I certainly have no idea myself.