But within Python 3, I think the deprecation system is decent. As far as I know, it really only occurs with "uncommon" APIs, so the vast majority of code will continue to work as expected on later Python 3 versions.
But within Python 3, I think the deprecation system is decent. As far as I know, it really only occurs with "uncommon" APIs, so the vast majority of code will continue to work as expected on later Python 3 versions.
I suspect you're referring to Unicode here. In that case, I think they could've just added a flag to Python 2's str type to indicate "is UTF-8" and deprecated the old unicode object. Then add some functions to extract code points or grapheme clusters or whatever else you need from the old school str object.
I might be in the minority, but I really like that Python 2's str could hold arbitrary binary data, of which UTF-8 is just one possibility. It had good interop with C, which I think is fundamental to a glue language like Python. I'd rather have fewer string types instead of more (one of my complaints about Rust too).
If you meant the print function, there were other ways to solve that too. The simplest might be to create a new name for the function and deprecate the print statement. So old code uses the statement "print 123" while new code is encouraged to call the function "echo(123)" or "ouput(123)". Bikeshed the actual name...
Note when I say "deprecate", I mean provide a timeline over several releases where it continues to work. Then issue deprecation warnings which can be silenced.
All of the newer features in Python 3 (@ operator, async/await, type annotations, etc...) could've been added in a mostly backwards compatible way. (Note: adding async wasn't really backwards compatible even in Python 3).
Anyways, hindsight is 20/20, but I really do think the path for Python 3 was a poor choice in comparison to other options.
No, and I explicitly mentioned UTF-8. My suggestion is that str holds arbitrary immutable binary data and that you have a method which can interrogate whether that binary data is valid UTF-8.
Yes, real world text is messy and there are lots of encodings, compression schemes, and exceptions (UTF-8 with byte order marks, overlong encodings, or surrogate pairs, as examples). If your main task is converting text between outdated or broken encodings, I don't have any problem saying you need a separate library and shouldn't burden the rest of the user base. Despite it's flaws, the majority of the world has settled on Unicode with a UTF-8 encoding.
"Special cases aren't special enough to break the rules."
Python 3's str can also hold arbitrary binary data. That ability was introduced in PEP 383.
My first thought as I started reading the PEP was, "Why did they bother adding the 'bytes' type if 'str' is just going to be able to hold everything anyways?"
After looking at more of it though, it seems like they're storing the binary octets as code points in one of several internal Unicode representations. Moreover, they're abusing (reusing?) the range of code points reserved for 16 bit surrogate pairs, but only using the low half of the pair. This is all clever in the bad way.
This seems like a real lack of taste to me, and I doubt the Guido from 1991 would've found it acceptable to have 'str', 'bytes', and 'bytearray' the way they are. (Let's ignore 'buffer' became 'memoryview' for now...) It used to be a simple and elegant language.
For that matter, why did the switch to unicode strings require breaking backwards compatibility? I don't understand why we couldn't just move to assuming all strings are utf8 unless specified otherwise. Keep [] byte oriented, add a .at() method to index by code point.
- conflating text with bytes, python had no way to tell whether given string is a text or bytes, because in 2.7 was the same thing.
- introducing unicode as an unicode type, this essentially made the problem worse because in addition to mixing text with strings, they added extra type to represent text, which was optional and some people used unicode some didn't.
- a cherry on top was implicit conversion, so if you passed unicode type where str was expected python implicitly converted it, and vice versa
This basically resulted in people writing a broken code in python 2, code that worked fine with us-ascii, but randomly blew up if there ware some non standard characters processed.
In python 3 instead applying another fix, they decided to do it correctly from the start. So str (Guido actually regretted he didn't called it text) is representing text and bytes are representing, well bytes. Python 3 also does not do implicit conversion as well.
I actually think they could perhaps help with migration, by adding to 2.7 one more import to __futures__ that disables implicit conversion and then treat bytes type as a distinctly different type than str. That could help people fix their code in 2.7 before migrating it, but anyway it's already too late to do it, also there is a hack that you could disable implicit conversion in python 2, but it showed that stdlib also relied on it heavily, so fixing that maybe was not worth it.
As an aside, there is currently a PEP that will allow for better concurrency __without__ removing the GIL, [PEP554](https://www.python.org/dev/peps/pep-0554/)
and a larger project overall for multi-core support, which is (hopefully) being included in 3.9 [Multi-core](https://github.com/ericsnowcurrently/multi-core-python)
Also, check out the SharedMemory objects[1] in 3.8. For my own use cases, this goes a long way toward making me no longer care about the GIL. I'd rather work on multiprocessing code than multithreaded code any time, and if you can easily share objects across processes, then the GIL isn't relevant to anything I'm working on.
[1] https://docs.python.org/3/library/multiprocessing.shared_mem...
If you want to do this in a more performant way you should be looking at a different language.
People talk about the GIL as if it’s just some drawback that could be removed and everything would be faster. There is a very good reason why it’s there.
It doesn't do that though. The GIL protects interpreter internals. It doesn't make Python code thread-safe, although it does make various operations atomic:
* native calls (which doesn't release the GIL) * individual bytecode instructions (which don't call into more python code)
The latter can make it seem like it magically protects your code, but even a simple method call is at least 3 instructions (LOAD_FAST, LOAD_METHOD and CALL_METHOD). Each instruction is atomic and the method call will be if it's a native call (e.g. dict.get) but that's it, the entire call itself is not atomic and you could thread-switch between LOAD_FAST and LOAD_METHOD, or LOAD_METHOD and CALL_METHOD.
It does not. What it does is create a bigger trap, because the code looks to be working in more testing situations, and will be more difficult to debug when it starts failing.
Threads may still be switched between instructions every sys.setswitchinterval milliseconds so every instruction boundary is still a potential thread switch point, leading to all the issues described by Glyph in his Unyielding essay: https://glyph.twistedmatrix.com/2014/02/unyielding.html
https://www.youtube.com/watch?v=P3AyI_u66Bw
See also Swift for another language whose implementation has performance issues caused by ubiquitous atomic reference counting.
There are other paths for improving performance that Python is embrancing: asyncio provides considerable performance gains to IO bound apps, and memory-sharing between processes allows for relatively cheap sidestep for those who find threads too slow for their use-case.
It's not Python's fault, but for example I was trying to get something running on Ubuntu 16.04 (Python 3.5) where the primary developer of it is running Ubuntu 18.04 (Python 3.6), and it seems like 3.5's version of the asyncio module has a bunch of bugs that make the tool in question basically unusable.
I know the usual advice— install newer Python from deadsnakes, use a virtualenv, etc. None of this stuff is the end of the world, but it can be jarring if you're only now getting off the Python 2.7 train and used to everything just working everywhere.
The process was not perfect, but in my book it’s a success.
I don't see how that changes what I wrote in a very meaningful way.
The problem is that every new Python version seems to break something. Python 3.8 was released almost two weeks ago, and some popular libraries still don't work with it, e. g. SciPy, Matplotlib, OpenCV, PyQt, PySide2 and probably a bunch of other libraries that depend on the above.
Code is a tool, treat it like any other. Just be aware of possible accumulating deferred costs. I don't know exactly what field your in, but just imagine any other piece of equipment or component going EOL, and the manufacturer only providing semi-drop in replacements. Or even if they are advertised as "drop in", you know that half the time they aren't, and something derpy happens on integration.