How can you be fooled by the U+202E trick? (2013)
galogetlatorre.blogspot.com
galogetlatorre.blogspot.com
With unicode, I really don't know what to do. What strings need to be sanitized (or validated)? File names? Urls? How do I sanitize them without causing agony for most of the world? Are there other unicode attack vectors?
The problem is honestly just that unicode is used to convey things like "this is a PDF" or "this is going to execute code" and it's attacker controlled. It's a terrible UX that's been abused for as long as paths have existed.
So I guess hope that operating systems will find a better UX for "this executes" and that they'll stop "a process executed" from being a game over situation for users.
Wink-wink ;)
They just turn your url into punycode, and more recently display that instead of the raw url: https://en.m.wikipedia.org/wiki/Punycode
Technically, punicode is the raw URL and old enough browsers will display all non-ASCII domains as punycode. It was only for a short while that browsers naively decoded punicode without restrictions.
Besides refusing to render U+202E when it is between characters of scripts that aren't bi-directional and flow in the same direction, the UI could display the string in ways that make it clear (e.g., different color, maybe add a warning tooltip, maybe add a dialog around an operation that could be dangerous, or maybe refuse to perform dangerous operations).
The general answer is to not allow (e.g., ignore it, remove it, but do not render it) in any identifiers, filenames, domainname labels, usernames, and other security-sensitive strings, U+202E between characters of scripts that are not bi-directional and which flow in the same direction. More generally, do not allow mixing of scripts in certain contexts -- for example, do not allow mixed scripts in a single domainname label. Some of these rules need to be implemented by, e.g., DNS registries, and may need to be tailored to their specific needs (e.g., you might find that .kr has to allow mixing of ASCII letters and Hangul characters because it's common in South Korea to add "ing" to names to make brands out of them).
UTR #36 doesn't go far enough, in my opinion.
[0] https://unicode.org/reports/tr36/
[1] https://www.unicode.org/reports/tr36/#Bidirectional_Text_SpoofingThe Trojan Source paper has a decent treatment of the area too, seen in a different application: https://trojansource.codes/trojan-source.pdf.
In the article it is shown how this can be misused to show executable files as a different file type:
file[U+202e]fdp.exe shown as fileexe.pdf
~~~~~~~~^^^^^^^
| |--- read this right-to-left, i.e. exe.pdf
|----------- this is the Unicode character Right-To-Left OverrideIf you "sanitize" Unicode you're just removing that functionality. Remove that character, and the right to left reversal wont' be there, right?
If you have to break it in the name of security and sanity, it shouldn't exist in the first place.
I am continually baffled about why the HN comment sections are such a hotbed of anti-unicode sentiment. Perhaps most of y'all don't have to write software that deals with non-English text in any significant way.
If you just didn't have buffers you wouldn't have to avoid buffer overflow attacks! Come on.
Yes, this aspect of SQL is poorly conceived because the only way is to manipulate it with text processing.
Parameters should be inserted into a query language using AST manipulation, like a Lisp backquote.
Now within a particular program, we can wrap something like that around SQL generation, and solve the escaping problems in one place.
That approach doesn't apply to character data, which has to be literal and primitive in many situations.
Some things are meant to be simple, and Unicode is massively complex. True, these are rare-ish corner cases, but every programmer knows it’s only ever a matter of time until you hit any possible corner case.
I’m not saying, screw people whose languages don’t fit into ASCII (I’m one of them btw), but I really don’t want to know the vagaries of Unicode representation to inspect a file name.
Maybe Unicode needs a “safe mode” where only outright characters are allowed? Or maybe the unavoidable complexity is a design flaw, for some uses anyway.
I feel like we’re close to someone discovering Unicode is Turing-complete...
> Maybe Unicode needs a “safe mode” where only outright characters are allowed?
What is safe for me may not be safe for you. I need RLM characters in my terminal emulator [1], you probably don't. And who knows, maybe RLM characters could prove to be an attack vector, by e.g. displaying filenames or other text backwards.Every Unicode character and code point exists because somebody, somewhere, needs it.
Every Unicode character and code point exists because somebody, somewhere, thought that somebody would find it useful.
Sure, just like they benefit from knowing the difference between 'V' and volts; 'J' and Joules; 'K' and Kelvins; 'A' and Amperes; π the letter, π the circle constant, and π the exotic particle; 'm', meters, and mass; 's', seconds, and position; 'g', grams, and gravity...
假設所有的意義都有自己的符號,那就太好啊
It's used by being hung on the wall, like a painting[1]. Like I said, it has no linguistic use, and thus it is not part of a writing system, which puts it outside the stated scope of Unicode. It is the exact equivalent, for weddings, of the upside-down 福 character that is hung for New Year's, or the wreath that Americans hang for Christmas.
But it does not correspond to anything in any language; there is no Chinese sentence whose spelling would include 囍. Note that the 新华 dictionary entry says "Double 喜. Generally used at happy occasions such as weddings.", and there are zero examples of the character being used. ( https://zidian.aies.cn/NDQ4MA==.htm )
The wording of the 新华 entry is almost identical with the beginning of the 汉语大词典 entry, and it's instructive to quote the rest of the entry:
> Character used at happy occasions. Commonly called "double 喜"[2]. Generally used at occasions such as weddings. Often cut from red paper (or gold leaf), or written on red paper, [then] pasted onto a door, window, or wall, in order to indicate a happy occasion.
[1] Actually, a much closer comparison would be the traditional magical talismans that use elements from the writing system in a freeform way to express various desired goals. https://en.wikipedia.org/wiki/Fulu
[2] This dictionary is nice enough to make the fact completely explicit that 双喜 is the name of the character rather than a definition. 汉语大词典 doesn't even attempt to provide a definition for this character.
Literally all emojis are not part of a writing system nor have proper linguistic use. Who on the earth write emoji by hand or even spoke it? Unicode already broaden the scope for so long, whether it is good or bad, nobody cares now. And beside that, a lot of symbols are also culture dependent. If you are not used to live at there, you just have no idea what it actually means.
Yes, that's true, and they've been a constant source of problems for Unicode ever since the decision was made to let them in. They stand in violation of Unicode's declared principles and purpose.
Note that 囍 is not considered an emoji by the Unicode standards, though that's what it is in fact.
> Who on the earth write emoji by hand or even spoke it?
But there is an example of this - the name of https://en.wikipedia.org/wiki/I_Heart_Huckabees [wikipedia title uses the word "heart", but the actual title uses the symbol] was frequently spoken aloud.
The real issue are the thousands of characters that appeared incidentally in some ancient text, either as typo or as a weird interpretation of a common character, which then ended up in the Kangxi dictionary, and then subsequently imported en-masse into Unicode.
Example, 𠒇 - the only known use (in non-ancient times) of this character was being the official name of 雷莊𠒇, winner of the 2017 Miss Hong Kong Pageant. According to her, she intended to write 雷莊兒 when she applied for her official documents, but somehow the officials interpreted it as 𠒇, which is really an archaic form of 兒 (at best). Reportedly she's changing her name back to 雷莊兒
Thousands of such characters exist, if you look at the page where 𠒇 is supposed to originate, more than half of this is obsolete -- https://www.kangxizidian.com/v1/?page=125#gv
(So yeah, you're correct in essence but picking on 囍 as an example probably doesn't really get your point across...)
囍 is used frequently, but it is not part of any writing system and does not convey any linguistic message. The opposite is true for 𠒇 - it is not used frequently, but it is part of a writing system and is used to convey linguistic messages.
To presume a character is "real" merely because it exists in the Kangxi dictionary is as valid reasoning as presuming a character is "real" because it exists in "some other encoding standard" that you've been dismissive about. It's just that Kangxi is the de-facto Han character encoding scheme before computer encodings were invented. (Unihan even contains all details about the radical, stroke and even page number where it was sourced from)
A lot of those characters appeared once in some ancient text, and took the Kangxi form due to transcriptions from scribes across the centuries, but we actually have no evidence that they are "real" (at any point in time). Some of these characters are known alternative forms of common characters, or are only known to appear in some ancient text before Han characters were standardized. Some are plain typographical errors. It's like a 3 year old child learning to write "ABC", which looks a bit weird, and then the unicode committee assigned 3 code points to them.
Those are entirely valid for Unicode. Han unification in Unicode is already considered a mistake. That's why "unified" code points now also have explicit, higher-numbered 'equivalent' code points that unambiguously refer to a particular graphical form. The graphical form is the whole point of Unicode.
I had to look them up so I might as well save others the time.
Not quite. That and because some twits are willing to cave in and add it, because adding crap to Unicode gives them a sense of purpose which makes them blind to the harm they are causing.
Operating system should already know if a file is executable and desktop should be using this information when presenting files instead of conventional names.
Yes it would mean that when you download a file you have to set the executable attribute before you can run it. It's already almost exactly like that on Linux. That's a good thing.
That already happens. If you set the "Details" view in Windows, it will display one file per row with various columns showing file metadata, including the file type. So for the example in the post, you'd see "[Word icon] annexe.doc - [the date] - Application" as opposed to "[Word icon] annexe.doc - [the date] - Microsoft Word 97 - 2003 Document".
In linux it's pretty obvious: https://www.rootusers.com/wp-content/uploads/2018/01/gnome-t...
The main advantage is that the UI for opening documents (vim file) is different from the UI for running executables (./file) which is different from the UI for running executables as administrator (sudo ./file).
This makes it impossible to accidentally execute a file that looks like a document.
cd ~/Téléchargements; chmod +x <try to remember the first letter><hit tab 15 times><miss it><go check inside the firefox download manager><restart><end up typing the whole name>
before being able to run it.The Windows UI could just render that character as a tofu box, and not evoke its reversal semantics.
Consider writing a bash script that makes a subdirectory with current date if it doesn't exist, downloads url from the command line argument into that directory, changes the permissions and runs it.
Then your workflow could be:
- keep terminal open
- Alt+tab to firefox, browse, copy link to executable file
- Alt+tab to terminal, ./handle.sh <paste the link> ENTER
- Alt+tab to firefox again
- rinse and repeat
There's many benefits - you don't get thousands of files in one directory (which can lead to performance degradation and even random errors), there's little risk that you run the wrong file, you don't waste as much time.
If it's from a GUI, file browsers have the capacity to that that bit too.
inotifywait -m -e create "$(xdg-user-dir DOWNLOAD)"|while read d e f; do chmod +x "$d/$f"; done
You can add "-r" if you want to also do it on subdirectories, and you can specify another directory instead of the default downloads directory.
Though I don't really recommend doing this since you should not be executing downloaded binaries often enough for this to be worth doing.
I literally do this multiple times per day. What in hell do you folk do with your computers?
Thanks for the script ! There's a few tens of thousands of files in my downloads folder, I hope this isn't going to pollute the inotify fds too much...
Hopefully not leaking company information to clever attackers.
I work with a number of clients in high security environments where you have to get permissions to perform operations like execution (or even setting chmod x on a file). It really does limit people running random things and causing destruction.
Executables tend to come compressed with a lot of other files nowadays, and compression software knows to set the executable bit when extracting them.
I have no idea in what situation I'd need to load so many executables via the browser.
If I download a pdf, I don't want it to run. I probably want to open it, but if I've gone through the trouble of downloading it onto the filesystem as opposed to just opening it in the web browser, there may be other things I want to do with it. Maybe I want to print it, or maybe I want to break it up into component parts. Or maybe I want to transfer it to a sandboxed environment before I open it because I don't trust it.
If you don't want to take an active role in maintaining security on your computer, then by all means use a commercial operating system. Pay a license and let Microsoft and Apple make those decisions for you. They have competent people who can program reasonable default settings for the average user. If you're running Linux or an open source operating system, you've already decided to take these responsibilities into your own hands.
That is a beyond ridiculous take. The last Windows OS that did not make me mad was Win98, and every time I have to use Mac machines I want to gouge my eyeballs out. I want an OS that lets me shoot myself in the foot as much as I want.
> If you're running Linux or an open source operating system, you've already decided to take these responsibilities into your own hands.
no, I just want to use the fastest possible OS for my use case on my hardware, which is Linux.
It's definitely an extreme view[0], and I'm open to being argued out of it, but I really think we should use a different and much simpler encoding for the vast majority of things even if it doesn't perfectly reproduce the written text of the language. Writing systems encoded the language in a way fit for the medium of paper, and maybe we need to consider computers a different medium and adapt our language to them differently too.
[0] and probably anglo-centric.
Most Unicode attack vectors seem to target people, they are ways to obscure from a person the true intention of what something is. As a developer it’s our job to try and mitigate this through how we display the data, warn about risks and block particular attacks. So it’s about working out how to prevent a user from making a mistake.
It seems to me in this case obviously Unicode became “a thing” after the development of a file extension. Clearly if it had been the other way around the file extension would have no meaning, maybe that’s what we need to move towards, store that metadata elsewhere on the file system.
Fundamentally Unicode from and “untrusted” origin shouldn’t be trusted in an executable context.
One could even translate the file extension into something more user-friendly (e.g. .exe -> Application), making it a "file type" column. Which is of course what Windows Explorer has been doing for decades already. Of course, that results in everyone turning the classic file extension back on...
The other problem with hiding the extension is that you then have "file name" collisions, two files will appear to have the same name but only differentiate by the "type". I think that's wrong.
> With unicode, I really don't know what to do. What strings need to be sanitized (or validated)?
Just yesterday I filed a bug on the KDE terminal emulator [1] for stripping away too many characters. I'm sure that the fine devs had no idea which characters are safe to leave in and which are not. I certainly have no idea myself.Eg. use my libu8ident. For technical details see the C++ proposal to use secure identifiers: https://rurban.github.io/libu8ident/doc/P2528R1.html
Unicode in URLs and usernames and filenames is just so easy to trick people with.
Will there be Unicode email addresses?
Let me rewrite that from the opposite perspective.
The reason I use it is so that I can write in my language.
Writing in a language the reader doesn't understand is ... not so fine.
Writing in a way to give the appearance of one message but the machine-recognised existence of another is ... wrong, malicious, and harmful.
At root the issue is that encodings and graphical presentations aren't the same thing. 7-bit ASCII is limited and constrained, but as a universally understood encoding those specific characteristics are useful benefits. Yes, it means that representations are limited. But that's the essential trade-off for a lack of ambiguity.
And even within ASCII, there are homoglyphs or near-homoglyphs: {0O,1lI, 5S} being the most frequently encountered. Kerning and ligatures may present others, as with {m, rn}. In historical documents, distinguishing {ſ, f}.
However, tangential to the grandparent's topic, I do imagine there are very specialized cases where Unicode is pretty objectively unnecessary.
well ok, it seems to me that it is more likely that in most cases Unicode probably has no effect on security one way or another, but there are a few edge cases where it does. Here is one case where using a particular character in a filename led to problems 8 years ago in an operating system historically known for not having the greatest security.
We can see here also a clear example of the edge case theory, almost every character in every language supported by Windows would not have caused any sort of problem when used in a filename. But there are a few characters where they would, this is probably an example of things programmers think they know about unicode or language or whatever, because we go around thinking that each unicode code point just represents a character in a language but some encodings have ways to represent little weird behaviors of particular languages, for example interlinear annotation characters https://www.unicode.org/charts/nameslist/n_FFF0.html and it is generally weird behaviors that end up being security hazards (not sure if there is any security hazard in interlinear annotation characters but wouldn't be surprised)
Apparently, yes. [0]
> making things more equitable for other languages and cultures
Well, yes. People often hate how compilcated Unicode is, but they also tend to forget that even ASCII was not sufficient for writing English in a proper way. I am not by any means old, but even I still remember the period on the web when non-Unicode encodings were relatively common and how problematic it was.
And, with assumption that by “other cultures and languages” you mean non-English speaking regions: Tolerate is a wrong word. Most people are not English native speakers. If we would go the route of deciding who is tolerating who, nations using Chinese Hanzi and its derivatives would be probably the first to claim they tolerate all the others by the speakers’ headcount alone.
Moreover, deciding what is/can be executed is not so obvious in my opinion. PDFs can have JavaScript embedded, SVG too. Flash games? Python scripts on computers with python installed? It is not just the names, but the very knowledge of what can be considered “executable.”
I know of both the over-strike method for traditional typewriters and extended ASCII. But I have no knowlede of a ASCII (extended or not) supporting every sign commonly used in English. Granted, it is not as significant as other languages missing whole letters, but signs like ”,“,—,–,’, which are often replaced with simpler ASCII alternative (",') or their combination (---,--).
Over-strike, from what I remember wasn’t implemented in any well-spread way with standalone ASCII, although I could be simply not aware of it. It gives us then a few more characters rarely used in English language in loanwords (née, naïve, façade).
Skipping diacritics with uppercase letters on displays with too little space happens with Polish language too, so I am aware of the practice.
In any case, yes, overstrike is just an approximation, but believe it or not, ASCII was designed to make it possible. And yes, it kinda sucks, and yes, it doesn't really make ASCII a multi-byte encoding, not exactly. But it's a funny thought that ASCII kinda almost was a multi-byte encoding. One can imagine ASCII evolving to treat BS<mark> as not unlike Unicode combining codepoints to make it possible to get proper diacritics even on capital letters, but ASCII would still have been a dead end, naturally.
The comment you're replying to just said ASCII without qualification. By definition this is not multibyte, or even whole byte (assuming by byte you mean octet). ASCII is a 7 bit standard, anything else is some other encoding.
I'm quite certain that ASCII was designed to make overstrike feasible. Overstrike just wasn't explicitly part of the standard. The quip about it being a "multi-byte encoding for Latin" was a joke made funnier (to me, and perhaps only to me) by having a kernel of truth in it.
This guy made a whole business out of it:
Are you overreacting? Unless you were proposing we rip it out, no, you're not.
Unicode isn't going anywhere. That's because the scripts that human languages use aren't going anywhere. And people do mix scripts, too.
There are answers though. First, Unicode has been a learning experience, and we're still all learning. Second, one of the outputs of all that learning that has happened so far is UTR #36 (https://www.unicode.org/reports/tr36/), which does cover a lot of these things.
I mean, obXkcd: https://xkcd.com/327/
But spelling out what you had in mind would be helpful here.
Still, a limited set.
Yes, errors are made and occur. They're reasonably easy to code defensively against.
Unicode ... vastly expands the attack interface.
> Unicode ... vastly expands the attack interface.
Limit yourself to 0..9 to be safe.
And how that compares to the number of 7-bit ASCII characters?
And how many special cases would have to be considered?
Scale matters.
The risks:reward ratio from 7-bit ASCII is low and manageable. The expressive capability is high. No, it's not a perfect representation for all languages. It is, however, a sufficient one, where common understanding is necessary.
At 128 vs. 10^20 possible codepoints, many in Unicode with side effects, the problem of deceptive or unintended use is exceedingly high in Unicode.
With far too many special cases to hold in human working memory, or even reasonably within most code-bases.
Same thing really.
So you don't know about mailoji.com? It was on HN a couple months ago I believe.
Well, that's a lot like saying "software, in general, is just a giant security flaw we tolerate because it does useful things". It's kinda the point aint it?
In 21st century, even Americans aren't able to write their own names and place names with pure ASCII anymore. So Unicode is there from necessity, not as an addon.
All technologies have intended / desireable, and unintended / undesireable effects.
The more powerful and flexible a technology, the more likely there are unintended / undesireable effects.
Many of those unintended / undesireable effects are themselves not obviously apparent, not immediately manifest, or both. All of which makes risk assessment all the more difficult.
Software, in that sense, is inherently a security flaw, as it virtually always manifests unintended and/or undesireable effects.
This isn't necessarily an argument to ban all software, though the Butlerian Jihad are taking notes. It is an argument, however, for acknowledging risks, raising awareness of them, and taking reasonable steps to guard against them.
And, yes, of course, email addresses in local languages, why the heck not? People in the world want to use their language! And yeah, even Hebrew, Tibetan, or Mongolian (in vertical script).
That comment of yours reads to me like an lazy post of a unilingual or uniscriptal cultural imperialist who really does not care about other languages.
Your attitude makes me rather angry, because Unicode is such an amazing achievement for the world, and your comment just comes off totally ignorant of that. It should be clear that combining the world's languages into a common standard is really, really difficult and will inevitably create something that is much more complex than your beloved ASCII. And it should also be clear that in such a complex standard, you cannot (a) solve all problems the first time you try, (b) you cannot try a second time, because no-one will adopt yet another such standard.
So just read those Unicode documents in order to understand. Those people really try and there are security consideration documents, and they are extended all the time. And take those new security warnings for what they are: they are problems with a complex system. It's expected. So when they get known, react calmly and figure out whether you need to fix anything.
C's flaws lead to the development of several different competing programming languages. Various flaws in cryptography lead to the development of algorithms that achieve the same result, but have far fewer implementation gotchas. IMO, the time has come for the development of a truly strict variant of Unicode that still supports the primary objective, but learns from Unicode's mistakes.
There are certainly contexts in which Unicode unambiguously and demonstrably leads to security weaknesses and issues. See generally homoglyph attacks.
At the heart of the lie and damage is the existence of a message which appears to say one thing but in fact says something different. It's the very limited nature of 7-bit ASCII, 128 characters in total, which provide its utility here. Yes, this means that texts in other languages must be represented by transliterations and approximations. That's ... simply a necessary trade-off.
We see this in other domains, in which for the purposes of reducing ambiguity and emphasizing clarity standardisation is adopted.
Internationally, air traffic control communications occur in English, and aircraft navigation uses feet (altitude) and nautical miles (dstance) units.
Through the early 20th century, the language of diplomacy was French. The language of much scientific discourse, particularly in physics, was German. And for the Catholic Church, Latin was abandoned for mass only in the 1960s.
Trading and maritime cultures tend to creat pidgin languages --- common amongst participants, but foreign to all, as distinguished from a creole, an amalgam language with native speakers.
A key problem with computers is that the encodings used to create visual glyphs and the glyphs themselves are two distinct entities, and there can be a tremendous amount of ambiguity and confusion over similarly-appearing characters. Or, in many cases, glyphs cannot be represented at all.
Where the full expressive value of language is required --- within texts, in descriptive fields, and in local or native contexts, I'm ... mostly ... open to Unicode (though it can still present problems).
Where what is foremost in functionality is broad and universal understanding, selectinga small standardised and widely-recognised characterset has tremendous value, and no amount of emotive shaming changes that fact.
As an example, OpenStreetMap generally represents local place names in local language and charactersets. This may preserve respect or integrity to the local culture. As a user of the map, however, not knowing that language or charcterset, it is utterly useless to me. Or, quite frankly, anyone not specifically literate in that language and writing system.
It's worth considering that the characterset and language in question are themselves, adoptions and impositions: English was brought into Britain by invaders, the alphabet used itself is Roman, based on Greek and originally Phoenecian glyphs. English has adopted or incorporated terms from a huge set of other languages (rendering its own internal consistency ... low ... and making it confusing to learn).
International communications and signage, at airports, on roadways, in public buildings, on electronic devices, aims at small message sets and consistent, widely-recognised symbols, shapes, fonts, and colours. That is a context in which the freedoms of unfettered Unicode adoption are in fact hazardous.
(Yes, many of those symbols now have Unicode code points. It is the symbol set and glyph set which is constrained in public usage.)
And the simple fact is that a widely recognised encoding system will most often reflect on some power structure or hierarchy, as that's how these encodings become known --- English, Roman Alphabet, French, German, Latin, etc. Small minor powers tend not to find their writing systems widely adopted (yes, there are exceptions: use of Greek within the Roman empire, Hindu numbering systems). Again, exceptions.
The problem is that nobody cared. Browsers invented punycode instead of following tr39, email ditto. But ok, at least something. Java did it, cperl did, rust did it.
Everybody else is vulnerable. Esp. most other programming languages, filesystems and login systems. https://github.com/rurban/libu8ident/blob/master/doc/c11.md
Now there will be no 'Documents', only 'documents'.
The comment I replied to already restricted us to ASCII letters, numbers, _, ., and -. I questioned why upper and lower case numbers should be allowed.
Calculating a hash or equality for a string always uses some kind of comparer logic. Being able to use a "raw" comparer would be one special case of that. In C#/Windows, you'd use
var cache = new Dictionary<string, Cached>(StringComparison.InvariantCultureIgnoreCase);
, or similar. These correctly calculate that "documents".GetHashCode() == "Documents".GetHashCode(), and that "documents".Equals("Documents"). You might think that this is more complex because of the case insensitivity, but it's only slightly so.
E.g. if you instead assumed case sensititivity and naively use a default here: var cache = new Dictionary<string, Cached>();
, then you'd actually be in MORE trouble because now the default comparison using locale-specific collation comes into play. So e.g. in germany files weiß,txt and weiss.txt would compare equal and thus also compare to the same hash (despite being two different files). A working linux lookup table with case sensitivity would look pretty similar to the first one var cache = new Dictionary<string, Cached>(StringComparison.InvariantCulture);->
(looks like) documentexe.pdf
With the \u202e character in, what does HN do with it?
documentfdp.exe