https://msdn.microsoft.com/en-us/library/aa365247(VS.85).asp...
They suggest avoiding <>:"/\|?* as well as all ASCII characters 0-31.
ASCII 0 can be really fun. Lots of filesystem APIs deal with NUL-terminated strings (like, all of POSIX) so a zero byte in the middle of your string just truncates it at that point. If you use something that tolerates zero bytes for your UI strings (like NSString on the Mac, maybe C++ UI frameworks dealing with std::string) then the full string may show in the UI and you just mysteriously get a filename that's shorter on disk than what you see on screen.
If you're in a position to enforce well-formed Unicode on all platforms, you're much better off. But many things (e.g. backup systems) don't get the option to just refuse files they don't like.
There is a very important takeaway of this: case-sensitivity. UNIX cannot be case-insensitive for file names because the mapping of lowercase to uppercase characters is dependent on the character encoding used, which it doesn't know. Windows can (and does) coalesce case for file names because it knows the character set in use and can consult the relevant mapping.
This difference in behavior produces all sorts of frustrating behavior when interacting between the two platforms, e.g. the classic case of Windows SMB mounting a share from a nix server that contains two files differentiated only by case. It'll show both entries but think they both point to the same thing. On the other hand, it's easy to create file names on a Windows device that are near impossible to name on nix. These are important things to be aware of if you ever implement a cross-platform network user environment.
Actually, it's not the character set in use. Windows uses a case mapping table which is part of the NTFS filesystem metadata. See for instance https://web.archive.org/web/20110308034840/http://blogs.msdn...
(Yes, this means that the mapping of lowercase to uppercase characters can change if the file is copied to another drive in the same machine!)
I kind of wonder if paths not being allowed to contain NUL or '/' was one reason why for codepoints that are represented through more than one byte in UTF-8 (-> all non ASCII codepoints) all bytes have the most significant bit set to 1 (https://en.wikipedia.org/wiki/UTF-8#Description) This makes it impossible to have multi-byte to contain valid ascii chars like `/`.
Note that macOS actually does decomposing unicode normalisation on file names, I guess because it makes handling case-insensitivity easier. (Just doing ascii case insensitivity also handles o+diaresis, but not the ö codepoint) https://developer.apple.com/library/mac/qa/qa1235/_index.htm...
/ and 0x00 for unix
:?"<>/|\* and chars 0x00 .. 0x31 for windows
'~!#$&%^; if there's a chance of filename being passed to shell w/o proper escaping.
Windows also forbids a bunch of filenames matching regex "CON|AUX|PRN|NUL|COM[1-9]|LPT[1-9]"
Also, ending filenames with space or period really messes up windows. File explorer can see it, but can't delete or rename it.
edit: fixed markup
Yeah, windows is kinda crazy inconsistent for some of these. I had a file (created under Linux) which ended in a space... drove windows nuts. Could list it, open it in some programs, but couldn't even open/rename by shortname under DOS or python.
As a related tip, if you need to name a file something like .foo in explorer, it rejects it as "not having a file name". But if you type .foo. then it accepts the name and silently strips the trailing period.
`echo missed one`disclaimer: i'm remembering something from the Windows 2003 era, so YMMV.
\/:*?"<>|
Surprisingly enough, FAR will deal with this 'somewhat' gracefully, but unsurprisingly, Windows Explorer will completely break.