Rust gets this almost right. Python gets this very wrong.
Rust gets this almost right. Python gets this very wrong.
Having worked in both, I'd say they both chose ideomatic solutions:
Rust: I can't prove this is utf-8, so if you want to use it as utf-8 you'll need to tell me what to do if it isn't.
Python: if you're in the common situation where everything is utf-8 and you want to just work with strings, go do the simple thing. Or you can be explicit about wanting to work with bytes, and that's good too.
(Though I think Python should throw an error instead of silently omitting non-utf8 files.)
>>> b'\xc0'.decode('utf-8')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xc0 in position 0: invalid start byte
>>> open(b'\xc0', 'w').write('foo')
3
>>> os.listdir()
['\udcc0']
>>> open(b'\xc0').read()
'foo'
>>> open('\udcc0').read()
'foo'[1]: https://peps.python.org/pep-0383/ [2]: cf. https://github.com/bup/bup/blob/master/DESIGN#L667-L729
You can actually reconfigure Python to throw an error instead of using a surrogate escape, but only (I think) changing something at compile time: https://docs.python.org/3/library/sys.html#sys.getfilesystem...