(Python's approach is that os.listdir("/") gives you a List[str] (silently omitting undecodable entries), while os.listdir(b"/") gives you a List[bytes]. That is, if you give the path as a utf-8 string it returns utf-8 strings, otherwise it returns bytes.)
The reality is that file names are "stringy" in nature--people expect to be able to do display them--and that means you need to have some up-front agreement on how to interpret those strings. In practice, on Unix systems, everyone has generally agreed that this is UTF-8, to the point that trying to not be UTF-8 generally causes interesting breaks in the system. It would be great if we could actually get the operating system to help enforce these rules, rather than placing the blame on other software for not correctly handling situations where the correct solution is itself incredibly ambiguous.
So I am afraid we can either a) indulge ourselves in wishful thinking, b) actively try to extinguish platforms that don't match our ideals, c) make an effort to be actually cross-platform, and not in "let's just build a tiny Linux model in a bottle for us to use and pretend the rest of the environment is not there" kind.
For untar'ing a tar archive: error out by default, but provide a flag or option to permit untarring using some sort of escaping to map the malformed names back into Unicode. I think here I'd map to something printable, though, like "\xnn" or something.
Paths might just be strings, but they aren't necessarily, and since Rust actually cares about types if you want strings you need to write the code to decide what to do about this, even if your "It's not UTF-8" case is just "Give up I can't be bothered".
Rust gets this almost right. Python gets this very wrong.
Having worked in both, I'd say they both chose ideomatic solutions:
Rust: I can't prove this is utf-8, so if you want to use it as utf-8 you'll need to tell me what to do if it isn't.
Python: if you're in the common situation where everything is utf-8 and you want to just work with strings, go do the simple thing. Or you can be explicit about wanting to work with bytes, and that's good too.
(Though I think Python should throw an error instead of silently omitting non-utf8 files.)
>>> b'\xc0'.decode('utf-8')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xc0 in position 0: invalid start byte
>>> open(b'\xc0', 'w').write('foo')
3
>>> os.listdir()
['\udcc0']
>>> open(b'\xc0').read()
'foo'
>>> open('\udcc0').read()
'foo'[1]: https://peps.python.org/pep-0383/ [2]: cf. https://github.com/bup/bup/blob/master/DESIGN#L667-L729
You can actually reconfigure Python to throw an error instead of using a surrogate escape, but only (I think) changing something at compile time: https://docs.python.org/3/library/sys.html#sys.getfilesystem...
That said, I actually didn't realize that os.listdir() silently omits un-decodable entries, which I find mildly alarming. This behavior isn't mentioned in the docs (https://docs.python.org/3/library/os.html#os.listdir) and seems out-of-character for Python, which usually raises an exception by default if data cannot be decoded to text.
Are you sure that this is actually what happens with non-decodable filenames? Reading here (https://docs.python.org/3/glossary.html#term-filesystem-enco...) and here (https://docs.python.org/3/c-api/init_config.html#c.PyConfig....), it suggests that encoding errors should be handled by surrogate escapes by default on non-Windows systems.
If it ever did that, it doesn't anymore.
Includes "Modify os.listdir(str) to ignore silently undecodable filenames, instead of returning them as bytes", but not later work where apparently this was changed to use surrogates.
I think that https://rushter.com/blog/python-strings-and-memory/ is a nice reference on that.