* unicode -> encode, u -> e (vowels)
* str -> decode, s -> d (consonants)
* unicode -> encode, u -> e (vowels)
* str -> decode, s -> d (consonants)
EDIT: This last line of the article pretty much sums it up: "We structured the transition thinking the community would come along with us in leaving Python 2 behind, but that turned out not to be the case and instead we have taken some more time and are using a Python 2/3 compatible subset of the language to manage the transition."
I wonder if python 3.6 or 3.7 will acquiesce to this and give an option for something like "from __past__ import str" like we have for __future__ in 2.7 to make these backward incompatibilities easier to deal with.
For example
old_function(bytes(myarg))
if it needs to be a literal:
old_function(b"literal")
That's what I do, and works for me so far.
As the past did not contain the __past__ that is not going to work.
There is one obvious way: strings are encoded to bytes, bytes are decoded to strings.
I think that reveals that the names really do have a problem. The problem is that "encode" sounds like "make this Unicode" to people who aren't familiar with Unicode.
Bytes hold a bunch of data in some encoding. It could be an image, UTF-8 or LZMA compressed ASCII. Once you know the encoding, to reconstruct the data you decode into a semantically meaningful form.
To put it another way, imagine the terms were "serialize" and "deserialize". Of course one serializes to and deserializes from binary data. Just replace "{,de}serialize" with "{en,de}code" and you're done.
Veedrac had a good analogy, think of text as something abstract, for example imagine text is an image or sound, if you want to store it in bytes you need to encode it, and to read back you decode it.
As to_bytes/from_bytes, actually python provides it too:
to_bytes -> bytes(<text>)
from_bytes -> str(<bytes>)