The hell that is filename encoding (2016)
beets.io
beets.io
1. Back in the days, we were using a Linux NFS server, with NFSv3, and out-of-the-box locale was iso-8859-1 (latin1). Life was good, except for occasional problems with people with strange non-latin1 names, or documents with non-latin1 names etc.
2. At some point, we switch to using UTF-8 by default. Telling users to use convmv to rename their files when they are ready to switch to the new defaults. Most people ignored this, of course, but files with now invalid utf-8 were mostly fine, just with the occasional "?"'s in the names.
3. Switch to NFSv4. Invisible to end users. NFSv4 per se requires that paths are UTF-8 encoded, but in practice the Linux NFS server and client just pass along a bag of bytes, so invalid UTF-8 just worked as fine as it did previously.
4. Switch from a Linux NFS server to a netapp.
5. User complains that files are missing. Initial comparison with the old Linux NFS server, which was still online, shows no problems. Problem occurs only on user workstation, not on admin box which has both the old Linux NFS and netapp directory trees mounted. Investigation on users workstation shows that in some cases lots of files appear to be missing, including ones which plain ASCII names.
- Turns out that the admin box had the netapp mounted with NFSv3, and thus everything appeared Ok there, including the rsync from Linux NFS -> netapp in the first place.
- However, when mounted using NFSv4, netapp follows the spec and does not like non-utf8 paths. Does it report an error then? Hell no, the NFS READDIR (READDIRPLUS?) message reply just stops returning directory entries when it hits the first one with invalid UTF-8. And thus you get a partial directory listing. GAAAH!
- So the solution was to run convmv centrally (from the admin box which had the netapp mounted with NFSv3) for the entire directory tree which had been moved.
The file names look OK after the copy on the Linux machine. However, when exporting the directory through Samba, the Macs Finder doesn't display files with accents in the names (though they appear correctly with "ls", weird...).
So the user copies the files again, using the Finder. Now I have files with exactly the same name (uhhhhh???):
# ls -l Mmo-1. -rw-rw-rw- 1 root root 8417218 6 sept. 2013 Mémo-1.aif -rwxr--r-- 1 test test 8417218 6 sept. 2013 Mémo-1.aif -rw-rw-rw- 1 root root 363175 6 sept. 2013 Mémo-1.m4a -rwxr--r-- 1 test test 363175 6 sept. 2013 Mémo-1.m4a
Yes, it looks like two files have exactly the same name, but actually they're different: one as "é" encoded as 0xCC81, and the other one (the "good one") as 0xC3A9. Why is that? Why does one work with the Finder, and the other doesn't? who knows.
Renaming the files to use NFKC normalization fixed it. In python, you could loop through the files and do something like:
os.rename(originalfilename, unicodedata.normalize('NFKC', originalfilename.decode('utf8')))
EDIT: You'll probably need to do this on a non-Mac system, linux for example should work.They were back when there were less than 2^16 characters in the Unicode standard. Back then each two-byte word in a filename corresponded exactly with a Unicode code point.
Now there are more than 2^16 but well under 2^32, Windows uses UTF-16 in filenames. That is, Unicode code points above 2^15 are obtained by a pair of special Unicode code points in the range 2^15-2^16 called surrogates; surrogate pairs need to be collapsed into a single code point when decoding the file name. Surrogates are exactly those things that Python uses on Linux to hide bytes that are not valid UTF-8. Here's the problem: it is possible to have unmatched surrogates in a file name (or in other places that Windows accepts UTF-16).
In summary, on Windows, you end up with effectively the same situation as Linux: file names that are supposed to be in one encoding (UTF16) but contain invalid data for that encoding.
For a fantastic read (from Victor Stinner himself) on all the work done and get how twisted this get when you want to be cross plateform and abstract it:
https://vstinner.github.io/python37-new-utf8-mode.html
Windows and Linux FS encoding, of course, are at the center of the challenge.
It's also the reason why we now have a __fspath__ protocol that allows any object to be converted to a file system path instead of having pathlib.Path inheriting from string.
NTFS (and Windows as a whole) does not use UTF-16, it uses UCS-2. It is a subtle difference, but surrogate pairs didn't exist in UCS-2.
The kernel generally uses the 16-bit equivalent of the old Pascal string, that being a 16-bit count of 16-bit pieces of UTF-16 data. This allows a 16-bit NUL to get into various places that make the Win32 API choke.
I usually call that ucs2-plus-surrogates, to make it clear that you may encounter unpaired surrogates, and thus invalid paths if you assume proper UTF-16.
Programs that internally use ShiftJIS, for instance, would stop functioning on UTF-16 enforced compatability or normalization. They're currently "broken" (as in, operating incorrectly) but in a way that works.
The reason they currently work is because bytes out == bytes in, so they can read the files they create, despite what mojibake the user sees.
You just described my music collection, aggregated over two decades and haphazard successive migrations. I've given up salvaging the corrupted names in an automated manner...
That is, on Windows, paths are fundamentally sequences of 16-bit words, just like on Unix paths are fundamentally sequences of 8-bit bytes. On neither system are paths fundamentally text.
NTFS is (usually) case preserving but not case-sensitive. So the OS needs to be able to tell whether EXAMPLE.TXT and example.txt are the "same" name, which means it needs case conversion.
Not everybody agrees about how this conversion should work. The most famous example is Turkish, but there are others. So there's an actual choice to make here.
If Windows baked this into the core OS, they might get pushback in countries where their (presumably American) defaults were culturally unacceptable.
If they made it configurable at the OS level, everything would seem fine until, say, a German tries to access a USB drive with files from a Turk on them and some files don't work correctly, or the disk just can't be mounted at all.
So, they have to bake it into each NTFS filesystem.
* http://www.edm2.com/index.php/Inside_the_High_Performance_Fi...
Has anyone tried this? Is it possible with FUSE? I would love to hear from people who know about this stuff - what are the obstacles? Or do you think FSs are fine the way they are?
The "open things by paths" thing, IIRC, is part of the reason Windows doesn't like to let you delete open files by default.
Making filesystems codeset-aware is not worth the trouble. It's best instead to just use UTF-8 locales everywhere. If you need to deal with other codesets, convert as needed, but don't use non-UTF-8 locales.
On ZFS with a fast ZIL fsync()/sync() function a lot like write barriers, which is what we really need. Actually, what we really need is for all filesystem operations to be available with async system calls, write barriers included.
Alone on windows, the maximum path length is 260 characters, except when you use extended-length paths which have a 4 character prefix and a maximum length of 32,767 characters. A sane API for reading files probably converts your paths to extended-length paths. But if you do the same thing for writing files your users start calling you insane again, because most of Windows (including Windows Explorer) can't open extended-length paths. So you would be creating files that only select software can even open, and which the user can't browse without third party software.
(An easier to ignore fun fact is that NTSF and Windows also support case sensitive names if you set the right flags in the file APIs. But nobody uses that, so it's probably save to ignore that (until somebody mounts EXT3 partitions in windows...))
iOS suggests that people are ok with this. /s
Honestly, normalization should probably always be set. It gets way more confusing than case-sensitivity (which can also be changed on ZFS!) already is.
Another way to put this is "shit in, shit out".
So: When receiving a filename that is not UTF-8, one could just emit a warning and ignore the file. That's what I would do if I wrote a music tagger, at least.
When someone else modifies a directory tree at the same time, we're bound to run into problems. That's just how it is. Actually there are synchronization facilities (e.g. flock()) for some of these types of problems. But these are seldom used because they lead to other problems.
I think they knew all that in the 70s, and just chose to be pragmatic about it.
They were reactions to the more structured access methods of the day.
The problem with invalid filenames is that doing anything at all with them might be impossible (without escaping of fixing). Your output (gui or terminal) most likely uses utf8, you can't pass invalid filename to it, so you can't even display it. Also in most situation you can't just ignore something. Imagine a text editor where user tries to open invalid file, how do you "ignore" that?
Many (if not most) application that process filenames do rely on those being valid text at least in some codepaths and they simply cannot work otherwise. The proper solution is to fail and let the user fix the problem.
Yep, that's what I meant. You need to do what is appropriate to the situation. Ignoring / warning are only two possible handling strategies. In many situations, straight-out failing is another valid one.
A file copy program should just not care and copy the darn thing. A text editor, basically the same, but maybe issue a warning.
Different tasks have different requirements, like the filename being text, the filename not containing spaces, etc. Given that these requirements are somewhat arbitrary, I think it's a fine choice to just not put any non-technical constraints in the guts.
There's an expression for it: Mechanism, not policy. You can always add policy on top, be it in the VFS layer, or as additional programs / classes of programs.
So many potential bugs would be easily fixed if the shell glob ('*') was more configurable.
Try saving a file or renaming a file to contain ":" using the GUI. It will not work. However, "/" is fine.
Try saving a file or renaming a file to contain "/" using the CLI. It will not work. However, ":" is fine.
A ":" in the CLI is translated to "/" in the GUI and vice versa. ":" was the directory separator in Mac OS 9 and earlier, which explains the behavior. The actual on-disk filenames will contain "/", which is translated to ":" for the POSIX API. This is for HFS and HFS+, not sure how APFS changes things.
So in the case of a Unix filesystem tree with various mount points, the FileSystem object would throw the exception iff the path you're trying to construct is illegal for the actual filesystem configuration in the system.
In many cases they offer the exact same semantics (damn you, mysql). Also they store data in files, and those files have to be named, so... back to square one.
I mean bad databases exist, sure, but the solution to that is to not use those.
> Also they store data in files, and those files have to be named, so... back to square one.
Not really - the database developers have handled the fsync, buffering and what-have-you for you, so you don't have to deal with them. (And FWIW serious databases generally offer the option of storing data on raw partitions).
* Windows and Mac filesystems are generally case-insensitive, so some users will have the file names in the playlist file in one case and the actual file names on disk in another format * Sometimes file paths cross between two different filesystems, because one is mounted in the other with a USB drive or over CIFS or similar. Sometimes these two different filesystems have different case sensitivities * There's no way to know how the playlist file was encoded * HFS+ normalizes file paths to Unicode NFD, but there's no guarantee that the paths in a playlist file will be normalized. Also, sometimes users generate an m3u file on a Windows system and expect it to just work on a Mac. Also, the filesystem nesting problem with network or USB mounts can happen this way too.
Ya know what kind of file names work virtually everywhere? ASCII ones.
The kind of file names that do work "virtually everywhere" are not ASCII, but rather are those who only use characters from the POSIX Portable Filename Character Set, which at 65 characters is just over half the size of ASCII (which has 128 characters).
* http://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1_...
This is why ZFS does form-insensitive directory lookups (and hashing)[0] rather than normalize-on-CREATE! I'm so glad ZFS got it right, and can stand as a model for all. (I implemented none of that functionality, though I code-reviewed some of it, specifically the u8_* functions in Solaris/Illumos, but I remember it took some doing to convince others that this was the correct approach.)
[0] https://cryptonector.com/2006/12/filesystem-i18n/ [1] https://cryptonector.com/2010/04/on-unicode-normalization-or...
IIUC the reason they did this is that they wanted directories to be canonically ordered on disk, and they thought decomposition would naturally yield better results than pre-composition. I'm not sure that's right, and frankly I don't care either, because the most important thing to note is that input methods (especially for European languages) by and large produce NFC, and most application software does no normalization at all, so disagreements as to form cause problems[0][1].
[0] https://cryptonector.com/2010/04/on-unicode-normalization-or... [1] https://cryptonector.com/2006/12/filesystem-i18n/
IMO it was a terrible mistake to normalize to NFD on create. Normalizing to NFC on create would still have been a mistake, but a lesser one.
”HFS uses 31-byte strings to store file names. HFS does not store any kind of script information with the file name to indicate how it should be interpreted. File names are compared and sorted using a routine that assumes a Roman script, wreaking havoc for names that use some other script (such as Japanese). Worse, this algorithm is buggy, even for Roman scripts. The Finder and other applications interpret the file name based on the script system in use at runtime.”
The bug (or part of it) was that some punctuation sorted before everything else.
> HFS Plus uses up to 255 Unicode characters to store file names. Allowing up to 255 characters makes it easier to have very descriptive names. Long names are especially useful when the name is computer-generated (such as Java class names).
I have to admit I let out a laugh at Apple's reference to Java class names...
- from wikipedia NTFS page [1]
So if you assume that NTFS filename is valid UTF-16 and convert it to UTF-8 there might be a problem. Basically they can be any sequence of 16-bit values.
[1] https://en.wikipedia.org/wiki/NTFSUnix and alike disallow NULs and /, for obvious reasons.
For example, compare these two in the command prompt:
start /Windows/Notepad.exe
start \Windows\Notepad.exe
I wouldn't blame this on poor parsing or other silly things though. I actually think it makes sense to use slashes for switches, because they are invalid filename characters, whereas dashes are valid, and hence get ambiguous (hence the need for "--" in *nix). I think it would've made more sense to disallow slashes as directory separators entirely, to avoid this for good.No...
assert( PathIsRoot(TEXT("C:\\")));
assert(!PathIsRoot(TEXT("C:/" )));
Also, I believe you meant directory separator, not path separator.Please don't be so tempted to take an antagonistic position and confidently declare other people wrong when you cannot possibly support your position in full... I see this very commonly on HN and I cannot tell you how extremely frustrating it is for those trying to help. It sucks away all the energy and enthusiasm we have for trying to help people get accurate information (meaning we might not even have the energy to bother to respond), and on top of that, you risk disseminating incorrect information. In this case, you simply could not have tried all the wide variety of Windows APIs, so at the very least, maybe say "in my experience" if something is only based on your experience.
For eg. in Rebol/Red - http://www.rebol.com/r3/docs/datatypes/file.html
A year ago I went on a trip and some combination of the humidity, the travel, and the 6 year old Thinkpad resulted in my laptop not booting.
I had been experimenting with Borg to backup the system, and so I tried using Borg to restore the latest copy onto the new laptop. Turns out that I have a bunch of files on my laptop that have names with weird characters in them: rips of my CD collection. I couldn't find any combination of settings and environment and locale that would allow borg to recover or skip these files and recover everything else.
Now, I had 2-3 other copies of the data (my pre-borg backups, the original SSD which was still readable, a few other rsync copies), so it wasn't a big deal.
But, as always, test your recoveries!
This is a big disconnect between "what most users expect" and "what systems actually do". Usually generally expect that filenames are sequences of characters - and today almost everyone expects that they must be in UTF-8 on a Unix-like system. That is not, of course, what most systems actually do.
TL;DR, basically, the lack of ability to tag strings in the system call API with codesets means that UTF-8 is the only plausible answer, and the ends (C library system call stubs, filesystems) have to apply whatever codeset conversions. But there's practically zero chance of C library system call stubs (and related functions) performing codeset conversions (can you imagine readdir(3) doing it?), which means that the only reasonable answer is to use UTF-8 locales and be done.
Even shorter: just use UTF-8 locales and be done.
chcp 65001
and bob's your uncle.I got the impression the problem in the comment above was too many extensions rather than too few. For example, you have use .docm etc rather than .docx if your document contains macros otherwise Word will refuse to open it (this is a security feature to prevent document viruses). But it sounds like there are many others.
The ASSOC .DOCX command shows an example of the first leg of the indirect mapping.
If the machine has a UTF-8 encoding (like, say, every modern system), it will try to treat filenames as valid UTF-8 strings and fail to back up files which don't fulfill that assumption. The "solution" is to run the TSM software with a single-byte locale like en_US.
I've seen a number of shops that were silently missing files from backup from old systems because of this problem.
... most do neither, but rather do ${complex thing emerging from combination of implementation details of runtime and backup tool, impossible to reproduce in any other runtime, likely platform- and environment dependent; the same backup likely restores in different ways on different machines, and the same source files create different backups on different machines; creating a backup on one machine and restoring it on another does not generally result in the same files; and I have not yet mentioned what might happen if you mount the same source file system from different platforms, because results might vary a lot; also, we are only talking about paths here, not any of the other plethora of things that can and will be different between any element in OSxFSxEnv}.
Sure it can. In this case, I'd say treating the filename as a bag of bytes is the correct way to go, as that's the way the OS treats them. Translating filenames between character sets should not be part of a backup systems job.
There are valid setups where different software on the same machine might be running with different character sets for legacy reasons. In that case there is no correct way to handle the filenames as text. But treating it as a bag-of-bytes will always work consistently.
Also, the one purpose of a backup system is to back up the files on the filesystem. If it can't back up some files that the OS considers valid, it's the backup software that failed.