The right way would have been utf-8, which is the only reasonable representation. Indexing could have been broken early in Python 3 or outsourced to a special type.
The right way would have been utf-8, which is the only reasonable representation. Indexing could have been broken early in Python 3 or outsourced to a special type.
The problem here is that file names aren't any kind of Unicode. You can have a utf-8 filename right next to a utf-16 file name right next to a Shift-JIS filename right next to one that's just binary gibberish that isn't a legal string in any encoding (that has any restrictions at all). They're just bytes with no indication of what coding scheme they are at all. You can't solve this problem with any Unicode solution. (Even the one that tries to wrap bad bytes in a recoverable way, whose name escapes me and I can't recall if it's actually a standard or not, would still make me nervous. You really need code that just treats the filenames as opaque bytes, and only later at display time tries to figure out if they're UTF-8 or not.)
It doesn't matter what Unicode model Python used. That may solve Python's other problems, sure. But filenames would always be some sort of problem. Even if you do just treat them as bytes, all that really does is make your potential errors deterministic, which is still better than a crash probably, but the filesystem's lack of encoding standards is a fundamental problem that can't be solved by higher code levels.
On Windows, paths are UTF-16 by convention, but also not forced to be valid. However, invalid UTF-16 can be faithfully converted to WTF-8 and converted back losslessly, so you can translate Windows paths to WTF-16 and everything Just Works™ [1].
There aren't any operating systems I'm aware of where paths are actually Shift-JIS by convention, so that seems like a non-issue. Using "UTF-8 by convention" strings works on all modern OSes.
[1] Ok, here's why the WTF-8 thing works so well. If we write WTF-16 for potentially invalid UTF-16 (just arbitrary sequences of 16-bit code units), then the mapping between WTF-16 and WTF-8 space is a bijection because it's losslessly round-trippable. But more importantly, this WTF-8/16 bijection is also a homomorphism with respect to pretty much any string operation you can think of. For example `utf16_concat(a, b) == utf8_concat(wtf8(a), wtf8(b))` for arbitrary UTF-16 strings a and b. Similar identities hold for other string operations like searching for substrings or splitting on specific strings.
Of course the implementation details are something that can and should be handled by a library instead of doing it manually.
On UNIX, paths are a sequence of bytes, with two bytes being sacred to the kernel (0x2F, used to separate path elements, and 0x00, used to terminate paths) and no other bytes being interpreted in any way. Any character encoding which respects the sacred bytes by not using them to encode any other characters is therefore usable to make UNIX paths; in fact, a UNIX path can contain multiple encodings, as long as they're all suitably respectful.
That requirement for respect means that UTF-16 and UCS-2 and UCS-4 are not suitable. UTF-7 is, however, as is UTF-8, and all of the ISO/IEC 8859 encodings are as well, not to mention a whole raft of non-standard "extended ASCII" character sets. In theory, UTF-16 in some suitably respectful encoding would work, too, but gouge my eyes out with a goddamned spoon.
I.e., their comment is an RFC "OUGHT TO", not an RFC "MUST".
Again, we're both aware that it's not guaranteed. But the convention these days is nonetheless UTF-8.
> mitts off unless the user explicitly tells you an encoding.
This doesn't adequately solve the problem, though. Typed languages have to emit some type of value: having that be the "string" type of the language is useful, as you can do things like printf a message that includes the filename. (Or display an "Open file…" dialog. Or…)
There are middle-grounds, such as byte smuggling and escaping. But if you take the stance that filenames are arbitrary bags of bytes (which is actually a subset: the reality is even worse) — then anything that returns a filename is stuck: it can't return a string ("mitts off"). You can take Rust's approach with Path (a type specific to Paths) but people wail about that all the time too ("why are there so many types?"), and you can't print it because it's not text!
"Error out if the file name isn't conventional" is a pragmatic tradeoff: bad file names will cause errors, but it makes basically all other operations much more tractable. It's not worth supporting insane file names.
There are workarounds, of course (such as just replacement-charactering "�" anything that can't be understood), and finding a format that can encode Paths when transmitting them, but these all take more time and effort. Allowing non-text files introduces unnecessary complexity and bugs into every single program that needs to deal with the file system.
Nonsense. Unix paths use the system locale by convention, and it's entirely normal for that to be Shift-JIS.
Not allowing strings to represent invalid Unicode is also a huge mistake (and essentially forced by the representation strategy that they adopted). It forces any programmer who wants to robustly handle potentially invalid string data to use byte vectors instead. Which is exactly what they did with OS paths, but that's far from the only place you can get invalid strings. You can get invalid strings almost anywhere! Worse, since it's incredibly inconvenient to work with byte vectors when you want to do stringlike stuff, no one does it unless forced to, so this design choice effectively guarantees that all Python code that works with strings will blow up if it encounters anything invalid—which is a very common occurrence.
If only there was a type that behaves like a string and supports all the handy string operations but which handles invalid data gracefully. Then you could write robust string code conveniently. But at that point, you should just make that the standard string type! This isn't hypothetical, it's exactly how Burnt Sushi's bstr type [1] works in Rust and how the standard String type works in Julia.
I have ranted for long hours go friends about the insanity of Python 3's text model before. It's mostly the blind leading the blind.
So, only build the index for random access if needed. Optimize "advance one glyph" and "back up one glyph" expressed as indexing, and you'll get most of the frequently used cases. Have the "index" functions that return a string index return an opaque type that's a byte index. Attempting to convert that to an integer forces creation of the string index.
This preserves the user visible semantics but keeps performance.
PyPy does something like this.
Like, the changes Python 3 made were honestly pretty subtle, and nearly 15 years later there's still people reluctant to upgrade. If they broke string indexing I'm pretty sure Python 3 adoption would make Perl 6 uptake look impressive.
Is this not the only thing you can really do? If strings could hold invalid unicode then they effectively become bytes and you now have to be wary of every possible string. I would rather Python just do away with strings entirely in the integration points with the OS and make you either keep them as opaque bytes or decode them.
> behaves like a string and supports all the handy string operations but which handles invalid data gracefully
Unless Guido was feeling particularly practical I can't imagine this ever making it into the stdlib because the choice of "gracefully" is somewhat arbitrary and application dependent.
First, why is Python unable to represent invalid path names as strings? Because internally it converts strings from UTF-8, UTF-16, or any other encoding, to a fixed-with array of decoded Unicode code points. The width of integer used to represent code points is determined by the largest code point in the string: if the string is ASCII, it can use a byte (uint8) per character; if the string is non-ASCII but all BMP, then it can use a uint16 per character; otherwise it has to use uint32 per character.
Why does Python do all this? So that you can have O(1) character indexing. If you gave up on that, you wouldn't need to convert the string at all, you could just leave it as (potentially invalid) UTF-8 data.
Suppose you get an invalid path on UNIX where paths are UTF-8 by convention? What does Python do with this string? It can't convert it to an array of code points because invalid UTF-8 doesn't correspond to a code point (well, it can if it's just illegal, not malformed, but in general, we have to consider completely malformed strings that don't even follow the basic UTF-8 format). So Python is stuck: it can only replace the invalid data with something like the Unicode replacement character. But then you can't do anything useful with that because it's not the correct name of the path you're trying to work with.
How does using UTF-8 to represent strings help? Because you can represent invalid strings: just leave them as-is and don't try to decode them unless you have to. Sure, you can't decode them as code points, but that's actually a pretty unusual thing to do. If someone asks for decoding, _then_ you can give an error. What about Windows where paths are UTF-16 by convention? You can convert them to WTF-8 and everything works out. (Described in way more detail here: https://news.ycombinator.com/item?id=33984308).
That’s not UTF8. That’s a bag’o bytes which might be UTF8. Very different thing.
> Sure, you can't decode them as code points, but that's actually a pretty unusual thing to do.
It’s not, any unicode-aware text processing does it implicitly. This means any such processing has to either perform its own validation that the input is valid, or it may fly off the rails entirely if fed nonsense. This also increases risks if security issues, either outright UBs, or the ability to smuggle payloads through overlong encoding.
True; I was careful not to call it that, but treating strings as UTF-8 by convention does make sense.
> It’s not, any unicode-aware text processing does it implicitly. This means any such things processing has to either perform its own validation that the input is valid, or it may fly off the rails entirely if fed nonsense.
In theory, but that's just not how most string operations actually work. If you have two UTF-8 strings and you want to concatenate them, you just concatenate the bytes. It would be ridiculously inefficient to decode the code points in each string and then re-encode them back into a destination buffer. If you have two UTF-8 strings and you want to see if one is a substring of the other and at what byte index, you just look for the bytes of one as a "substring" of the bytes of the other. Again, it would be ridiculously inefficient to decode the code points in each and do matching on code points. But what if the strings aren't valid UTF-8?! Both of those operations work just fine even if the strings aren't valid and produce sensible, intuitive results.
If you're implementing a browser or a terminal that has to actually display UTF-8 as characters then sure, you have to actually decode characters. Similarly, if you're parsing text somehow, then you have to decode characters. But many program only do concatenation and search and other operations like that which are actually implemented in terms of byte sequences, not characters.
You specifically called it UTF8, repeatedly. The very comment I quoted asserts that "Utf8 would help deal with the issue [of garbage inputs]" (in its denial of the opposite assertion). You also did it in https://news.ycombinator.com/item?id=33986421
> If you have two UTF-8 strings and you want to concatenate them, you just concatenate the bytes.
That's not a unicode-aware operation, it's mostly a unicode-irrelevant operation (though unicode awareness can be useful in edge cases because of special grapheme clusters, but that's very task-specific).
> But what if the strings aren't valid UTF-8?! Both of those operations work just fine even if the strings aren't valid and produce sensible, intuitive results.
If your content is not actually UTF-8, you can end up with UTF-8, thus changing the semantics of the content. You can also end up with overlong UTF-8, which also changes the semantics of the content in a worse way.
The comment I originally quoted was yours. The second quote is adapted from a statement you denied. It is thus your statement. Let me put both sections together since you seem unwilling to do so:
>> Utf8 would not help with the issue in the article in any way.
> It's not at all obvious how it helps, but it does.
So you are stating, unambiguously, that "Utf8 would help deal with the issue [of garbage inputs]".
> In the comment you link to says "UTF-8 by convention".
The comment I link says:
> Treating paths as UTF-8 works very well
Which is either
1. wrong
or
2. nonsensical, given the later statement that you should not "require your UTF-8 strings to be valid", which would make them not UTF-8
> If you're concatenating two strings that are both invalid UTF-8, there's not much you can do that's better than just concatenating the bytes together... which is exactly what treating them as byte arrays would end up doing
But that's the point innit? You're asserting semantics which don't hold and which you break with no regard.
> (but it's less convenient).
Is it now? Here's the concatenation of two strings:
a + b
here's the concatenation of two byte arrays: a + b
You're right, the inconvenience makes me shudder. What horror. What indignity.Ok but I’m not saying to do that. I’m saying if you have not-utf8 strings don’t call them UTF8.
> How do you want to write a program that opens a path who's name is invalid UTF-8?
That’s not my problem given I’m not advocating for that.
Seems like "giving up" would've been better choice, considering just how rare operation that is. Or alternatively doing the conversion lazily the first time operation needing runes instead of bytes happen.
Most string operations are not accessing string by index and most of them even at O(n) would be fast enough because n is small. Like in typical "get a file name, extract some info from it", you're doing extraction once and anything after that doesn't need character indexing, because you already got the relevant data.
How is that better than just handling paths as `bytes`?
LOL this is great, so true