The choice to have this in a language with the safe/ unsafe distinction works very nicely because in so many languages you'd have this promised UTF-8 type and then in practice everybody and their dog uses the unsafe assumed conversion because it's easier, but in Rust you're pulled up short because that conversion needs an unsafe block, your local style may require (and good practice certainly does) that you address this with a safety comment explaining why it's OK and... it just isn't, in most cases. So you write the safe conversion instead unless you really need not to. This is a really nice nudge, you can do the Wrong Thing™, but it's just easier not to.
MFC had a similar problem for years. It had a CString class which was ANSI or Unicode, depending on how your compiled your app, but moderately often you needed the other one, so it should have had a CStringA and CStringW too, with nice conversions between them.
In most of these file format cases what you've got is &[u8] or &[u16] and maybe it's the NonZero variant instead, so I think it's fine to be explicit that's what is going on and maybe in the process remind you to check - is this data UTF16LE? UTF16 with a BOM? UCS2 with a nod and a wink? Just arbitrary 16-bit integers and good luck?
But like I said, I favoured "picking nits" long before I learned Rust, so mileage may vary.
Since references can't have destructors (they don't own the data like an OsString does), it means that the standard library can't give you a newly-allocated string without leaking it. Since obviously it isn't going to do that, the &OsStr must instead just act as a view into the underlying &str. And the conversion can't enforce any extra restrictions on the input string without breaking backward compatibility.
The overall effect is that whatever format OsStr uses, it has to be a superset of UTF-8.
[0] https://doc.rust-lang.org/1.77.2/std/ffi/struct.OsStr.html#i...
That said, even now in 2024, it's not clear how much of a bet Windows is making on UTF-8 versus UTF-16.
Then when 32 bit came to be we had the other variations on top.
Got to love leaky abstractions.
Reference to “Celestial Emporium of Benevolent Knowledge”?
As a meta-joke I was also considering:
APL (j/k)
It's a problem with Unix filenames where the encoding is just a convention. A Python program that doesn't take great care can crash on a parameter that takes a filename even if that filename is just passed to a function like 'open' so no sanitisation or conversion is necessary.
This is a very real problem in hivex, our Windows registry library, where the Python bindings don't really work well. The Windows registry is a hodge podge of random encodings, essentially whatever the program that wrote the registry key thought it was using at the time. When parsed through Python as a string this means you'll get unicode decoding errors all over the place.
Also more in this article: https://changelog.complete.org/archives/9938-the-python-unic...
Don't get me wrong it's still painful and annoying and bug prone, but the point is, it's encodings, it's always going to suck no matter what.
UNIX filenames are NOT necessarily printable text, they're byte strings. Don't treat them as printable text. They're sequences of bytes not containing 0x00 or 0x2F, but with no encoding.
Same for Windows registry keys. Don't mistake byte strings for text.
Text is a byte string with an encoding which describes which byte values are valid and which characters each byte (or sequence of bytes) corresponds to. If the encoding information is discarded, it stops being text and becomes just a byte string.
But they nearly always also happen to be printable text. Treating them as printable text is soooo convenient 99.9% of the time.
Having them be almost printable text, but not quite, is an API design wart that is begging to be used incorrectly. It's the inverse of an affordance.
The POSIX filename rules are a "here be dragons" sign. If you can avoid dealing with them, it's best to do so. If you can't avoid it, you'll need to parse your inputs into valid text, and fall back to a byte string handling path (or just fail with error) if they're not text. You can't safely just assume they're text & treat them as such, it doesn't end well.
(IIRC, at least one of the BSDs has actually moved to forcing filenames to be UTF-8 and refusing path names that aren't UTF-8. Would only that Linux moved down that path as well so that we could be done with this farce.)
I'd agree it'd be nice if we could restrict filenames to valid UTF-8. But that's not the API that existing filesystems provide, nor what (most of) the existing OSes enforce.
Threading both of these needles at once basically requires viewing paths as potentially-invalid encoded text.
How else would you decode a string without knowing its encoding? You can either guess (and risk invalid result/decode errors) or store this information somewhere. This is universal and true in every language. In most cases today people choose to guess utf-8.
>>> 'abc' + 'def'[1]
abce
>>> b'abc' + 'def'[1]
TypeError: can't concat int to bytes
It's been responsible for a lot of bugs in my code. They copied the "bytes is an array of integers" thing from Java. Big mistake. Python is not Java.Them treating filenames as strings is another bug factory on Linux. Almost no one unit tests filenames with invalid encodings, so the result is a whole pile of python3 programs will fall when given perfectly valid input, whereas those same programs in python2 were fine.
It's a very odd outcome given python harks from Linux.