There are also limits to how far the UTF-8 illusion can go on Windows: while on Unix and friends a path is fundamentally a 0x00-terminated, 0x2F-separated sequence of 8-bit quantities[1], on NT a path is fundamentally a(n unterminated) 0x005C-separated sequence of 16-bit quantities, and Win32 puts a varying number of layers of makeup[2] on that. Thus on Unix you must be prepared to handle invalid UTF-8 in a filename, but can expect to roundtrip any byte sequence (sans 0x00 and 0x2F), and on UTF-8 Win32 you must be prepared to handle arbitrary WTF-8[3] and cannot expect to roundtrip any byte sequence (isolated surrogates can merge, though I don’t know if UTF-8 Win32 is willing to accept such invalid WTF-8).
Note that Rust does not use the UTF-8 interfaces on Windows (neither does it use the fundamental UNICODE_STRING APIs, however).
[1] https://yarchive.net/comp/linux/case_insensitive_filenames.h... (of course, Linux filesystems have since developed case-insensitive mount options)
[2] https://googleprojectzero.blogspot.com/2016/02/the-definitiv...
Plus, you don't need to be prepared to handle invalid utf-8 in filenames on unix, the fopen call can just be made to fail if needed.
As being prepared for invalid UTF-8 on Unix, well, it depends. If you refuse to run with non-UTF-8 LC_CTYPE or to accept invalid UTF-8 in user-provided file names, I suppose that’s on you. (Though I sure hope you are not writing an implementation of rm or tar!) If you’re trying to erase or move everything in a directory, though, you’ll have to either deal with whatever’s there or at least recognize that the action may fail.
It's not all roses on unix or rust side either. In unix filenames are not utf-8, which leads rust having fun things like OsString.
(Python's approach is that os.listdir("/") gives you a List[str] (silently omitting undecodable entries), while os.listdir(b"/") gives you a List[bytes]. That is, if you give the path as a utf-8 string it returns utf-8 strings, otherwise it returns bytes.)
Paths might just be strings, but they aren't necessarily, and since Rust actually cares about types if you want strings you need to write the code to decide what to do about this, even if your "It's not UTF-8" case is just "Give up I can't be bothered".
That said, I actually didn't realize that os.listdir() silently omits un-decodable entries, which I find mildly alarming. This behavior isn't mentioned in the docs (https://docs.python.org/3/library/os.html#os.listdir) and seems out-of-character for Python, which usually raises an exception by default if data cannot be decoded to text.
Are you sure that this is actually what happens with non-decodable filenames? Reading here (https://docs.python.org/3/glossary.html#term-filesystem-enco...) and here (https://docs.python.org/3/c-api/init_config.html#c.PyConfig....), it suggests that encoding errors should be handled by surrogate escapes by default on non-Windows systems.
I think that https://rushter.com/blog/python-strings-and-memory/ is a nice reference on that.
If it ever did that, it doesn't anymore.
Includes "Modify os.listdir(str) to ignore silently undecodable filenames, instead of returning them as bytes", but not later work where apparently this was changed to use surrogates.
The reality is that file names are "stringy" in nature--people expect to be able to do display them--and that means you need to have some up-front agreement on how to interpret those strings. In practice, on Unix systems, everyone has generally agreed that this is UTF-8, to the point that trying to not be UTF-8 generally causes interesting breaks in the system. It would be great if we could actually get the operating system to help enforce these rules, rather than placing the blame on other software for not correctly handling situations where the correct solution is itself incredibly ambiguous.
So I am afraid we can either a) indulge ourselves in wishful thinking, b) actively try to extinguish platforms that don't match our ideals, c) make an effort to be actually cross-platform, and not in "let's just build a tiny Linux model in a bottle for us to use and pretend the rest of the environment is not there" kind.
For untar'ing a tar archive: error out by default, but provide a flag or option to permit untarring using some sort of escaping to map the malformed names back into Unicode. I think here I'd map to something printable, though, like "\xnn" or something.
Rust gets this almost right. Python gets this very wrong.
Having worked in both, I'd say they both chose ideomatic solutions:
Rust: I can't prove this is utf-8, so if you want to use it as utf-8 you'll need to tell me what to do if it isn't.
Python: if you're in the common situation where everything is utf-8 and you want to just work with strings, go do the simple thing. Or you can be explicit about wanting to work with bytes, and that's good too.
(Though I think Python should throw an error instead of silently omitting non-utf8 files.)
>>> b'\xc0'.decode('utf-8')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xc0 in position 0: invalid start byte
>>> open(b'\xc0', 'w').write('foo')
3
>>> os.listdir()
['\udcc0']
>>> open(b'\xc0').read()
'foo'
>>> open('\udcc0').read()
'foo'[1]: https://peps.python.org/pep-0383/ [2]: cf. https://github.com/bup/bup/blob/master/DESIGN#L667-L729
You can actually reconfigure Python to throw an error instead of using a surrogate escape, but only (I think) changing something at compile time: https://docs.python.org/3/library/sys.html#sys.getfilesystem...