I think there would have been much less of a problem if encode and decode were far more obvious, unambiguous and intuitive to use. Probably without there being two functions.
Still a problem of course today.
I think there would have been much less of a problem if encode and decode were far more obvious, unambiguous and intuitive to use. Probably without there being two functions.
Still a problem of course today.
Here's how you remember it: "Unicode" is not an encoding. It never was, it never will be. Of course, the data must be encoded in memory somehow, but in Python 3, you cannot be sure what encoding that is because it's not really exposed to the user. From what I understand, there are different encodings that string objects will use, transparently, in order to save memory!
You always "encode" something into bytes, and "decode" bytes back into something. There should be exactly two functions, because the functions have different types: "encode" is str -> bytes, "decode" is bytes -> str. Explicit is better than implicit.
output = input.decode(self.coding)
With Python 3, I instantly know that "input" is bytes and "output" is str. In [1]: len("नि")
Out[1]: 2
So close, but it's not really a string or bytes, but a codepoint array. You can even iterate it: In [2]: for c in "नि":
...: print("Character {}".format(c))
...:
Character न
Character ि[1] Sadly surrogate escapes break this. Rust's division into String:OsString:Vec<u8> makes that aspect a lot cleaner.
Strings have to be arrays. It is inevitable. We are on Von Neumann machines with L1 cache lines that are something like 64 bytes wide, so making strings into something other than arrays is a complete non-starter. Because strings are arrays, we use array indexes to slice strings. The indexes are integers, because that just makes sense on a Von Neumann machine.
So your complaint is that you can slice a string in "bad" ways. What makes it bad? You're trying to prevent people from slicing up grapheme clusters, which is a noble goal. But in practice, we get the indexes from methods like str.find() or re.match(), and we treat them as more or less opaque, so we don't end up slicing grapheme clusters very often in practice. Grapheme cluster slicing requires a big Unicode table in memory anyway, so in the rare case that you need it, you can pay the high performance cost for using it. In the meantime, formats like JSON and XML are defined in terms of code points, so using code point arrays eliminates a class of bugs where you could accidentally make malformed XML or JSON, which would then get completely REJECTED by the receiving side, causing your Jabber client to quit or your web browser to show a bunch of mojibake.
And let me ask you this: what do I get when I write:
x = "a"
y = "\u0301"
Is the resulting len(x + y) == 1? But len(x) == 1 and len(y) == 1? Can you tell me what the correct behavior is? Might you end up with bugs in programs because len(x + y) != len(x) + len(y)? Or do you introduce extra code points into the string when you concatenate them?Please, tell me what you actually think is correct behavior. It is far, far more useful than pointing out that something is "wrong" on purely semantic arguments.
https://www.mikeash.com/pyblog/friday-qa-2015-11-06-why-is-s...
What you can drop is the idea of the index being an integer with any particular semantic meaning. Rust uses byte indexes with an assertion that the index lies on a codepoint boundary, for instance.
Outside of the ASCII range? I doubt it. That most likely means you are doing something fundamentally wrong.
EDIT: Or a JSON parser, etc.
It seems like this whole mess comes from conflating bytes and characters. They are not the same, any more than integers and booleans are the same just because you can jnz in most assembly languages.
Why on earth not? Any slice of a string (with a non-pathological internal encoding) with ends on codepoint boundaries is itself a valid string. Losing that information sounds like a whole lot of pointless hassle and potential for harm. It also forces the internal string encoding to be part of the interface, since taking a slice of a string and getting a byte array requires knowing how the string is encoded to use it, but getting another string means that such details are private.
It's true that they're byte offsets, but that doesn't matter. The parser doesn't care whether it's on the 10543th grapheme or the 10597th. It just wants to know where its stuff is.
I mean, if one has a hashmap
hash = {"foo": 1, "bar": 2}
we say hash["foo"] is an indexing operation, and "foo" is an index. Seems the same to me.There are plenty of uses for getting the length of a string where you will not be running into combining characters, clusters or any other things which would require special handling.
(remember that on the internet, all the data you receive is initially in string form, and very often we want to receive data representing fixed-length values such as phone numbers or credit card numbers, so "you're wrong for wanting to know the length" is both a non-starter and, well, just plain wrong)
How do you know that? Any user supplied text may contain those things.
>remember that on the internet, all the data you receive is initially in string form
It's not. It's bytes.
>very often we want to receive data representing fixed-length values such as phone numbers or credit card numbers
This is a good example of why strings shouldn't have a length property - people don't understand what length actually does in most languages. If you think that a len() function can verify that the contents of a text field is in the format of a CC number then you don't understand the len function. It can't do what you are asking. See the examples given in the GP (len("\u0301") for example).
The fact that the people who advocate for length functions want to use them for things that they are broken for is exactly why length functions shouldn't exist.
Which is why, in addition to things like length checks, we use other checks. But length is a quick and easy check for many types of values, and should be supported.
I've been down in the depths of Unicode behavior a time or two. I know what lurks down there, and I know from experience that your "strings should never allow these operations" stance is just as dumb as the more-common "I don't know what Unicode really is" stance. Can length checks and slicing and other operations run into trouble in some corners of Unicode? Sure. But if "this might cause trouble" were grounds for forbidding programmers ever to do something, we wouldn't have computers, any useful software at all, the internet or most of the other things we like and take for granted despite the fact that we occasionally have to fix bugs in it.
Now, get down off your snobby horse and join the rest of the real world, where we know that something might not be guaranteed reliable in all logically-possible cases, but find ways to work with it anyway because in 99.99% (or more) of the cases our code deals with it's good enough.
You say "wrong" as if that's a quote from me. But it's not. Yes, strings have to be arrays, but how they work depends on the semantics of your language.
I'm not entirely happy with Python 3 because I don't think they really took their time to improve on the warts. I doubt there's a one-size-fits-all solution.
But that's a choice and not the only one. Sure you're likely to implement your Unicode string as an array of some type. But that doesn't mean that the only sensible approach is to expose that array directly to the user. Or more specifically as the primary interface.
For my money, Swift has probably the most comprehensive Unicode string API I've seen [0]. What they do is, essentially, have an opaque String type but support various "views" as properties. The main view is called .characters and represents a collection of grapheme clusters. They also have properties that present Unicode scalars, utf8 encoding etc.
Their API is complex, no question. But then proper handling of Unicode is complex. But it does show that there are other options than simply exposing the Unicode scalars.
BTW to your point on len(x) + len(y). The answer is 2 if you define len() on Unicode scalars but 1 if you define it on characters. Why? because len("\u301") should be 0. It's not a character, it's a base modifier. It is, of course, true that getting this right is likely to be substantially more expensive than getting it wrong but that doesn't mean it can't be done.
[0] https://developer.apple.com/library/ios/documentation/Swift/...
You learnt ASCII. Then you learnt about escaping to get around limitations in ASCII. Then you learnt that Unicode got around non-Latin issues with ASCII.
Each step has added cognitive load. It seems surprising to me that the next step wouldn't.
I'm not saying that the string API in Python is perfection, but it had a specific role to fill in the larger Python ecosystem, and it does that very well.
Anyway... I would say that len(x) == 0 should be the same as saying that x is empty. According to the Unicode standard, the number of grapheme clusters in "\u0301" is 1, anyway! And on most systems, the string will display as an acute accent in its own space, and you can place the cursor before or after it.
Not like the set of valid "cursor positions" in a string isn't a locale-dependent value, anyway. Grapheme clusters can be thought of to be locale-independent but they are NOT the same as cursor positions.
If I recall correctly the modifier on it's own isn't correct. There is a base that should go before it to produce a valid character. I don't recall the base unfortunately and I don't have the spec to hand.
I mean this from the standard: "Nonspacing combining marks used by the Unicode Standard may be exhibited in apparent isolation by applying them to U+00A0 no-break space. "
You can have valid Unicode that doesn't do this. But it's not intended to represent the character as some kind of textual thing. It's intended so that you can build up Unicode strings sensibly. Same idea as allowing mis-matched surrogates.
Perl 6 has long specified an interesting and rich approach to Unicode. Have you explored it?
(An initial implementation for an initial release of the Rakudo compiler was worked on over the last few years and declared done, modulo bugs, a few weeks ago. A recent blog post provides a friendly overview.[1].)
One notable difference is that Swift adopts an iteration-only view of strings for dealing with characters whereas Perl 6 does not. One consequence is that indexing operations in Perl 6 are O(1) (time) at the expense of potentially (and for the current implementation, typically) greater use of RAM.
> BTW to your point on len(x) + len(y). The answer is 2 if you define len() on Unicode scalars but 1 if you define it on characters. Why? because len("\u301") should be 0.
Are you 100% sure of this of your last sentence ("should be 0")? I get that it's intuitively reasonable that an isolated base modifier be length 0 but I thought the Unicode standard specified (or at least recommended) that it be 1, and I note that Perl 6 treats isolated base modifiers (or at least this particular one) as length 1:
> ("a").chars 1
> (0x301.chr).chars 1
> ("a" ~ 0x301.chr).chars 1
So both Swift and Perl 6 correctly report the length of "a" as 1 and an "a" concatenated with "\u301" as 1 but disagree on what to do with an isolated base modifier.
> It is, of course, true that getting this right is likely to be substantially more expensive than getting it wrong but that doesn't mean it can't be done.
Indeed.
[1] https://perl6advent.wordpress.com/2015/12/07/day-7-unicode-p...
String searching & matching return an opaque String.Index type. I assume that under the hood they do Boyer-Moore on bytes and String.Index is a byte offset, but the important thing is that String.Index values are not convertible to integers, and so you never run into the case where a user passes in a byte offset that would slice a grapheme in half. Instead, String.Index has properties for the next and previous index and a method to advance by an arbitrary amount, so you'd access the 7th character after a dash in a string as myString[myString.rangeOfString("-").advanceBy(7)]
Swift gives up on the idea that len(x+y) = len(x) + len(y). This is just something you have to remember; in UI programming, however, it makes a lot of sense because adding an accent to an existing string isn't going to make it take up more width in the text box. (x + y).utf8.count == x.utf8.count + y.utf8.count, however.
https://developer.apple.com/library/ios/documentation/Swift/...
If I were to adapt this to a domain like Python's, where string processing is pretty important, I'd allow indexing & slicing by integers (including negative integers), but I'd define them in terms of String.Index types, which iterate over extended graphemes under the hood:
str[2] == str[str.start.advanceBy(2)]
str[-3] == str[str.end.advanceBy(-3)]
str[str.find('foo') + 3] == str[str.find('foo').advanceBy(3)]
str[0:-3] == str.substring(str.start, str.end.advanceBy(-3))
del str[0:3] == str.removeRange(str.start, str.start.advanceBy(3))So do fullwidth characters have a length of 2?
If not, how is this useful for finding the size of text in a text box anyway?
(0) Dependent types: The type of valid indices into a string depends, well, on the string. It's this type, not `uint`, that should be used by functions like `str.find()` and `re.match()`. This makes it a compile error to, say, slice a string with an index to another string. Internally, or course, all indices are represented as `uint`s, but your program has no business knowing this. It's an abstract type. IMO, this solution is particularly useful for systems languages, where the increased flexibility and fine-grained control are worth the increase in complexity.
(1) Coalgebras: `str.find()` shouldn't return an index. Instead, it should return in O(1) time, and allocating O(1) additional space, the left and right slices of the original string, split at the desired index. (If the string can be split in multiple places, `str.find()` should return an iterator whose element type is such left-right pairs.) For this to work, it's fundamental that the left and right slices not be allocated in their own buffers, but instead share the original string's buffer. IMO, this solution is best for high-level languages, where simplifying the API is well worth a little loss in flexibility.
Neither of these APIs allows the programmer to split strings at the wrong places.
Perhaps this is because Unicode is an exhaustive standard. :)
Seriously, I hear ya.
> every argument had the same hole that I see in your post: What, exactly, is a "character"? Please answer me that
I'd say the answer varies depending on who's talking, what they're talking about, and who's listening.
If it's a programmer interested in listening to the Unicode consortium and interested in what Unicode.org documents define for "what a user thinks of as a character" then the answer is, according to those documents, a "grapheme".
> then you can write a proposal for how the API should work, exactly
Right.
It can get, and has indeed gotten, better than that; one can write specifications, build reference implementations, and try things out in battle for a few years. Several text processing systems (including some programming languages) have gone this route and their results should be taken in to account. (ICU is perhaps the go to reference implementation.)
> and why this works with Indic and German and Korean and everything else.
For designs and implementations of text processing systems that claim an aspiration of progressing toward fully following the Unicode specification, the simplest answer to "why does it work?" (or "why it is expected to work") with Indic and German and Korean and everything else is of course the Unicode standard itself.
For those of us who aren't blessed with the patience of a saint and thus haven't read the entire spec from start to finish a few times, one has to rely to a degree on those who do, even if it's really hard to follow what the heck they're saying.
> You're trying to prevent people from slicing up grapheme clusters, which is a noble goal.
Given that grapheme clusters correspond to what the Unicode standard specifies for identifying and preserving the integrity of "what a user thinks of as a character", it seems less a noble goal than a fundamental eventual goal for any text processing system that explicitly claims to aspire to support the world's human languages via Unicode compliance.
> we don't end up slicing grapheme clusters very often in practice.
Which "we" are you referring to?
The western world has dominated text processing to date does. This must not mean that "we", by which I mean all humans, would best continue practices that marginalize those with native languages not well served. (By ignoring the central importance of "what a user thinks of as a character"; again, I'm quoting Unicode consortium documentation.)
Or, more parochially, what about the "we" that's just programmers who want to be able to increase their confidence that they are properly processing Unicode compliant text?
> In the meantime, formats like JSON and XML are defined in terms of code points ...
Code points are the middle level of Unicode, sandwiched between byte encodings at the bottom level and graphemes at the top. There's nothing wrong with a text processing system stopping at the middle level provided folk don't lose sight of the fact that it stops at the middle level.
Thinking that code points are "what a user thinks of as a character" mostly works out OK when processing text that's English or the like, but is definitively not Unicode compliant to the degree it's used too broadly.
> Can you tell me what the correct behavior is?
I would suggest that the most appropriate authority is the Unicode consortium and its Unicode standard.
> Might you end up with bugs in programs because len(x + y) != len(x) + len(y)?
Oh for sure. But it's part of the Unicode standard. It's one of several apparently strange discontinuities that the Unicode effort introduced last century in an attempt to deal with the combination of human language reality and computer system realities. They may have chosen the wrong sweetspot. But now Unicode is a significant part of reality. I, for one, don't think it's wise or even realistic to abandon the Unicode standard as it is...
https://github.com/python/cpython/blob/master/Objects/unicod...
Clever is relative. The Python 3 unicode string objects makes very little sense given real world situations. It wastes enormous amount of memory in many common scenarios for no reason at all other than to facilitate O(1) access to charpoints which is completely inappropriate for text processing anyways.
Worse than that: it can blow up by another 50% the moment someone encodings it to utf-8 as the encoded version is cached around.
The object is too clever for its own good.
>>> import sys
>>> def charsize(x):
... return sys.getsizeof(x + x) - sys.getsizeof(x)
...
>>> charsize(b'A')
1
>>> charsize('A')
1
>>> charsize('\xff')
1
>>> charsize('\uffff')
2
>>> charsize('\U0010FFFF')
4Except it does not help. When you render out a template it starts out as latin1, then the first time it hits a unicode char the whole thing re-encodes into ucs2 and then when it sees an emoji it re-encodes to ucs4. In practical terms you will see UCS2 and UCS4 all the time.
IMNSHO, most modern languages should be storing strings as UTF-8 and give up on random access by characters. You almost never need it; in the most frequent case where you do (using indexOf or equivalent to search for a substring, and then breaking on it), you can solve the problem by returning a type-safe iterator or index object that contains a byte offset under the hood, and then slicing on that. Go, Rust, and Swift have all gone this route.
This has nothing to do with (inefficient) encodings, and their bad historic namings, probably derived from perl.
Of course any name like byte2utf8 or just to_byte or b2u8 would have been better than encode/decode.
And of course cache size matters nowadays much more than immediate substring access for utf8, so nobody should use ucs-2 or even ucs-4 internally. This is easily benchmarkable.
That's not true as of 3.3. It will store the compact representation, which in most cases is ASCII, which is UTF-8 compatible. https://www.python.org/dev/peps/pep-0393/
So unless you go outside of UCS-1, you don't need to reencode anything.
Only a few standard member functions of string (find, rfind, index, rindex) return integer subscripts into a string. Those could return an opaque index object which has an index and a link to the string. Such objects could be used for string operations which need a string index. If an index object is forcibly converted to an integer, the array of indices has to be generated, but otherwise, it is unnecessary. You could even allow adding or subtracting integers from an index object, which would walk the UTF-8 string forwards or backwards as indicated.
This would make UTF-8 strings usable with subscript and slice notation without exploding them to wide characters.
I'm not sure "encoding" tells us that the output is bytes. Surely information can be encoded in other ways. Names like .to_bytes() and .from_bytes() might have saved a lot of trips to the documentation.
Information - yes. Strings... I'm not sure. What other representations are you thinking about?
When save some "things" I read from utf16 file and want to save them to Postgresql who's client encoding is set to utf8, what do I do?
It's important to point out that "Unicode" in Windows is an encoding. It means specifically UTF-16 LE, many people like me learned the concept "Unicode" on Windows got confused to hell when learning Python.
Encoding and decoding are pretty well defined (though, I've never thought about the formal definition before). When you have an entity in its native form, it needs to be _encoded_ for the purposes of communication (in a broad sense). The encoded message can then be to be _decoded_ back to the natural form. There is no ambiguity.
Really, the reason people get them mixed up in Python is because Python 2 totally stacked it by adding str.encode and unicode.decode.
In Python 2, you can _decode_ unicode to unicode – which it does by silently _encoding_ as ascii first. This operation is total madness.
>>> u = u'\xe9'
>>> u.decode('utf8')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
UnicodeEncodeError: 'ascii' codec can't encode character u'\xe9' in position 0: ordinal not in range(128)
The error here is in the step where you _encode_ the unicode to acsii, something you didn't ask it to do at all.And similarly, you can encode str to str (where str is really bytes in Python 2, another issue that adds to the confusion).
Don't get me wrong. I feel your pain. When I was using Python 2 I also got confused about what form things were in and where I needed to encode / decode.
Honestly, once I switched to Python 3, that cognitive overhead just totally vanished. str is the natural form of text, and if I need to store / communicate it I _encode_ it to bytes (utf8, generally). When I'm loading a stored/transmitted message, I _decode_ it to its natural text form.
There are edge cases that make certain situations more complex, but in terms of general usage, I feel Python 3 really got this stuff right.
The confusions is around people viewing byte strings implicitly as ascii codepoints, and not fully understanding what the things they're looking at actually are.
output = ''.join(str(x) for x in some_list)
Can you spot the danger here? Can you imagine how many hidden str() calls in every lib/modules?That's maybe because JSON is by the standard supposed to be utf-8, always.
The point of the comment is that everyone knows that you json encode native entities to the utf8 representation, and json decode from utf8 to native entities.
The same is true of unicode. It's not ambiguous at all.
Only utf-8 is a very specific encoding, so converting to it is straightforward, whereas unicode is an umbrella term and the most ambiguous thing you can imagine.
Even saying "the same is true of unicode" shows confusion, as utf-8 is itself unicode. As are numerous different encodings, and then you have normalization and other subtleties.
If some is "in Unicode" it's a list of codepoints - that is the "thing" to be encoded. I don't see that ambiguity. It's a list of numbers, it's not fluffy kittens in the network or something. UTF-8 is a way of transcribing that list of numbers as 1s and 0s. That is the encoding process. Once it's encoded to 1s and 0s, it's not anything. It's just 1s and 0s.
One person's utf8 encoded 1s and 0s may be another person's audio signal. It doesn't really make sense to talk about the encoded data as being utf8. It's the encoded / decoding process that's utf8. Sure, it might be utf8 decodable.
In terms of json, it's the same deal. You have the "thing", simple nested objects/lists/numbers in memory. It only makes sense to talk about json when you're talking about the process of transcribing those.
In both cases you have a "thing" (I think of it as "real", but in a way it's an abstract concept - how does Python represent a dict, I have no idea), you can encode those things into a tangible list of 1s and 0s.
That's why they're the same. And people aren't generally confused about the json encoding decoding process.
Conversely, you can decode the byte message into something readable, ie. your Unicode string.
* unicode -> encode, u -> e (vowels)
* str -> decode, s -> d (consonants)
There is one obvious way: strings are encoded to bytes, bytes are decoded to strings.
I think that reveals that the names really do have a problem. The problem is that "encode" sounds like "make this Unicode" to people who aren't familiar with Unicode.
Bytes hold a bunch of data in some encoding. It could be an image, UTF-8 or LZMA compressed ASCII. Once you know the encoding, to reconstruct the data you decode into a semantically meaningful form.
To put it another way, imagine the terms were "serialize" and "deserialize". Of course one serializes to and deserializes from binary data. Just replace "{,de}serialize" with "{en,de}code" and you're done.
Veedrac had a good analogy, think of text as something abstract, for example imagine text is an image or sound, if you want to store it in bytes you need to encode it, and to read back you decode it.
As to_bytes/from_bytes, actually python provides it too:
to_bytes -> bytes(<text>)
from_bytes -> str(<bytes>)
EDIT: This last line of the article pretty much sums it up: "We structured the transition thinking the community would come along with us in leaving Python 2 behind, but that turned out not to be the case and instead we have taken some more time and are using a Python 2/3 compatible subset of the language to manage the transition."
I wonder if python 3.6 or 3.7 will acquiesce to this and give an option for something like "from __past__ import str" like we have for __future__ in 2.7 to make these backward incompatibilities easier to deal with.
For example
old_function(bytes(myarg))
if it needs to be a literal:
old_function(b"literal")
That's what I do, and works for me so far.
As the past did not contain the __past__ that is not going to work.
First you need to think of text and binary (bytes) as two distinct types.
The text most of the times is how you interact with user, so things that you will display to the user, or what user enters to you.
Now, the text is stored in unicode (how, it's not our concern, Python abstracts it from us), the bytes is representation how it's stored in files, send over a network etc.
Now if you need to store text in a file, or send over network you need to encode it (most of the time as UTF-8), if you receive data which supposed to be text you decode it. It might be helpful to think of a text as a sound and bytes as mp3. You encode sound as mp3 to store it in a file, and when you want to play it back you decode it.
There are some things can cause a confusion. For example if you write text to a file, python will automatically apply the conversion. Things like that are there so you don't have to do the conversion every single time.
I'll take an occasional 2 seconds in the REPL over the headache of debugging codec issues any day of the week.
They are not. What is unintuitive is the default encoding in Python where an encode can trigger an implicit encode and the other way round. The `encode()`/`decode()` availability of strings was never a problem has you have many bytes -> bytes and str -> str codecs;
But then function would do something completely different. "\x01\x02".encode('zlib') for instance is a bytes to bytes operation. The problem is that "foo".encode('utf-8') does not give you an exception. If the coercion would not be enabled you would get an error:
>>> reload(sys).setdefaultencoding('undefined')
>>> 'foo'.encode('utf-8')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "/usr/local/Cellar/python/.../encodings/undefined.py", line 22, in decode
raise UnicodeError("undefined encoding")
UnicodeError: undefined encoding
That's not any worse than what Python 3 does: >>> b'foo'.encode('utf-8')
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
AttributeError: 'bytes' object has no attribute 'encode'I think the inclusion of things like zlib (or rot13 or whatever) was a conceptual error that just fosters confusion.
We should not optimize languages for idiots. There is nothing confusing about such an operation for anyone who can use their brain. Python 3 still contains those operations but instead of x.encode(y) you now do encode(x, y).
Of course this is a real thing and it's very frequently used. Yes I am in favour of it as the codec system in Python is precisely the place where things like this should live.
>>> "foo".encode('zlib')
LookupError: 'zlib' is not a text encoding; use codecs.encode() to handle arbitrary codecsYou can use bytes(string,encoding) to replace encode(). Unfortunately it doesn't have a default encoding, which makes it a pain to use. And str(bstring) isn't symmetric, it can't replace decode().