This cements the CPython team as sort of a known "don't waste your time there" in PL research communities, distancing themselves from the people most likely and willing to help maintain their runtime.
This is by a core python dev who is currently employed by Microsoft to enact this plan.
In the last roughly 10 years, Google, Instagram, Dropbox, Facebook, and reddit all invested a lot of time into trying to make Python fast and was pushed away with harsh no's. There are dozens of forks and branches with basic optimizations that were rejected.
I wish everyone the best of luck. But there's a side of me that's disappointed since this somewhat shows that it's only when these ideas come from inside the core maintainer team that they get acceptance. That's going to make it difficult to grow the team in the future.
https://news.ycombinator.com/item?id=8772124
It is amazing that someone who has been unfriendly (sometimes in a nasty political manner) to so many people still has a cult following him and is ostensibly a CoC proponent, now that the CoC is used to protect the inner circle.
If I were an expert assigned to this project at Microsoft, I'd try to get out of it as soon as possible. CPython is too political.
Is that true?
The talk you link was about a very secretive fork of Python that Instagram was very adamant about not open sourcing until very recently (in fact it's not clear that what's open sourced is their full project).
Other than that until recently all I've ever seen is either breaking the C-API or no clear commitment from the maintainers that this was a sustained effort to support and not just a drive-by complicated patch.
In more recent times we have cinder, pyjion, pyston, and others. And CPython Devs are up-streaming either their work or their general approach.
What goes in CPython depends on whether Guido and his seconds likes it or not. And if it is their ideas and their implementations it is treated differently than if it is from outsiders. Then sometimes they suddenly change their minds and something that previously weren't ever going to happen gets implemented. JITting which, apparently, now is on their road map is one such example. Works for them and Python continues to be a great language, but frustrating for contributors.
I think we've seen recent languages that Strings are basically one of the hardest things to "get right". Just look at Swift and Rusts multiple attempts to implement strings even though the authors all had many years of experience of working with string implementations in other mature languages.
Python doesn't have the luxury of messing up strings, it was the number 1 reason that migrating from Python 2 to 3 took a decade. Unlike lower level compiled languages it can't stick in multiple string attempts and tell the users to pick one. So Core Devs are on the hook for maintaining any implementation for the life time of Python.
But maybe this example really was a concerted effort by someone willing to maintain Python strings for as long as needed.
A large part of why Python 3.6 is a better port target than 3.0 is because they slowly and silently stepped away from their very strong opinions about "what is text", and "what is bytes", and what each of those things meant. Even back when the model was debated, it was clear that the operations the core maintainers cared about (manipulation of individual codepoints [0]) was a dead-end, as the Unicode Standards committee was just about realizing that codepoints weren't an ideal unit, and grapheme clusters were added to the spec. (my dates might be a bit off here, it's all a blur I've mostly forgotten)
Python 3.0 was released as basically a completely nonfunctional piece of software. You basically couldn't write anything that touched a binary file in it. The built-in zipfile module, I remember, crashed if you tried to store or load a binary file stored inside a .zip, the only tests at the time stored .txt files, and the bug wasn't fixed for a long time. I remember others having troubles with the email and mime modules, though I didn't have any code that worked with those, personally.
I had left the community by that point, but I have heard tales of the mountains Armin Ronacher had to move in order to add the u'' prefix back to Python 3; the core team was against it, instead believing that if 2to3 wasn't working for people, it should be fixed. Alice Bevan–McGregor spent years making a version of the WSGI specification that worked with Python 3's text model. We honestly could have shaved 5-6 years off of that disaster if they had just listened to the community; pretty much every opinion we had shared eventually came true, and the Python 3 text model these days is in a very different, and much healthier, place.
[0] One of their major strongholds was that they really needed O(1) indexing of codepoints (it wasn't clear why this was a priority, and I remember arguing that wanting this is a sign of poorly written code). I think in Python 3.8 they finally caved on this, with a PEP and an implementation that can use UTF-16 or UTF-8 internally.
It's pretty clear that there were a lot of mistakes in the Python 2 to 3 transition, I'm certainly not trying to defend every choices of the developers.
As for the rest, I'll see if I can dig up some logs, but a lot of the discussion I was involved in took place on IRC, and my logs were lost a few server wipes ago. You might find some old posts of mine on python-ideas.
I've been planning on making a fuller blog post on why I think Python's text model is the wrong one, but Manish has a very good post [1] which covers most of the reason. Basically, Unicode code points don't net you anything over bytes other than O(1) indexing, which is completely useless, and in fact, is in many ways a worse representation. There is no operation where having a list of Unicode code points helps you more than having a list of bytes; well, one that isn't completely broken in practice.
More drastically, the team wanting "Unicode everywhere" meant forcing things that clearly weren't Unicode into Unicode. File paths have not been valid Unicode on any popular system. Windows uses its "UTF-16" which is really more UCS-2 and allows unpaired surrogates, and POSIX says node names are bytes but cannot use ASCII NUL or /. These limitations were hit days into the prototyping process on real-world use cases, which is why there's an algorithm to shove arbitrary binary data into unpaired surrogate characters (see PEP-383 [2]), to be used on file paths. This should have blew a huge hole in the whole scheme, but they pushed onwards. The utopian future of "Unicode everywhere" was more important than dealing with the practical reality of existing systems.
Similar things happened with console output, where it's now assumed that LANG will be set correctly. But no, when I ssh into a server in Japan that's set to use EUC-JP, Python will output EUC-JP-encoded bytes, which gets shovelled over ssh as bytes, which the terminal emulator on my laptop misinterprets as UTF-8. Yes, I know ssh is supposed to tunnel certain LANG envvars as well. No, I don't remember why it didn't in this case. But a lot of people were hitting this, that in Python 3.7 they finally added UTF-8 mode (PEP-450 [3]). There's still a lot of people for whom this general "Unicode console" idea breaks [4] [5].
Pretty much every part of the standard library was riddled with bugs when they rolled this out (as mentioned, I hit bugs in zipfile, others found bugs in email/mime). For a long time, the only way to know if handed a file-like object is going to be in bytes or not was isinstance(file.read(0), bytes), which is generally ugly, so a lot of modules didn't support bytes even when they should have.
Practically everything after has been walking back this text model in favor of something where bytes and str are not even that different anymore. Python 3.5 brought with it PEP-461 [5], which added % formatting back to bytes (after a very long and very tiring discussion with the core maintainers [6]).
I could go on and find more PEPs, but I'm done for now.
[0] https://www.python.org/dev/peps/pep-0393/
[1] https://manishearth.github.io/blog/2017/01/14/stop-ascribing...
[2] https://www.python.org/dev/peps/pep-0383/
[3] https://www.python.org/dev/peps/pep-0540/
[4] On POSIX system, https://stackoverflow.com/questions/11741574/how-to-print-ut...
[5] On Windows, https://stackoverflow.com/questions/17918746/print-unicode-s...
Can you please expand on this? I thought the problems with Python's new text model were due to backwards incompatibiltiy.
Unfortunately, while the developers removed the implicit encode/decode in Python 3, they also made the gulf between str and unicode (now called bytes and str) much larger, removing large swaths of useful functionality from bytes, and doubling down on the Unicode code points nature of the new unicode type, despite it being very apparent by then that this was not a good text model to base a language on anymore.
Backwards compatibility was a large issue indeed, but IMO they broke backwards compatibility all to introduce a subpar stricter text model that delivered on far fewer promises than tney were hoping, and in my opinion is worse than Python 2's text model. A more reasonable approach can be found in other languages, which allow iterating over sequences of byte strings by iterating on the fly. Swift, Rust and Go all get this more correct.
Where 'just about' means 1996 at the latest, the release of Unicode 2.0 (cf chapter 5 [1], p. 21, Character Boundaries).
What are you referring to? Rust's str type is the only one I know about [1], and it's really good as a general purpose string representation. Were there other string implementations from before the 1.0 release?
[1] There's also String, which is just the heap-allocated version of str, and OsStr/OsString which represents the operating system's string type for file system paths.
Id love to see the core team develop some humility and realize their engineering is barely even mediocre and accept patches from well meaning people that know what theyre doing
It was like pulling teeth - he didn't know what property testing was ("some haskell thing?" - his words) and so he wouldn't update it ("are you implying that how Dropbox writes tests today isn't good enough?" - his words), so I couldn't write generated tests.
Sort of ironically, property testing was used elsewhere in the company to great success, and in fact we were all told the scary story of the 1 hour where empty passwords let anyone log into any account - a bug that a proeprty test is very well designed for.
I walked away pretty embarrassed for him and decided I wasn't interested in further interaction.
I also wouldn't call it overkill. It's trivial to do generated tests.
And rightfully so. If a function is long enough to need multiple lines, it's also long enough to be named.
Complaints about this seem to come from the functional language circles, where chaining 10 .filter()'s, .map()'s, etc. together in a single line is considered not only normal, but desirable and idiomatic. Why bother naming the intermediate results, right?
The ability to stuff a million operations on a single line that, for whatever reason, seems so coveted by many is quite clearly one of the least important aspects of language design. Vertical space is cheap and plentiful. Human brain capacity is tiny. Therefore, optimize for readability not clever one-liners.
This is the big benefit of a BDFL, having single a person with good taste be able to veto bad ideas. Guido, for the most part, has good taste so Python's ended up quite well designed.
To go full circle, pandas, an extremely popular python library, forces you to do exactly that, because the straightforward imperative loop is slow as a glacier.
Good taste would be having proper lambdas. Python is a very inconsistent language and among the last things I am reminded of when I think "good taste".
As a user of other languages, but only occasionally of Python, this type of reasoning (arbitrary limits because we think if you exceed them you’re probably structuring your code wrong) is totally alien to me. Why should the language have opinions about how I should or shouldn’t structure my code?
Imagine if Rust or C++ had an arbitrary rule that you couldn’t have more than 10 functions in a class/impl; maybe it’d usually be correct style but people would still think the language making it a hard restriction was ridiculous.
Isn't that what attracts so many people to Python in the first place - that it does have strong opinions, and often enforces them on how people should structure their code.
Python, in many places, optimizes for readability, rather than flexibility, and there is a community of users that really appreciate that. Not to say the feeling is universal (clearly isn't) - but it's at least an answer to the "why?"
I seriously doubt it. What attracts people to Python is that:
1. It has a big standard library and until the last ~10 years installing 3rd party packages was beyond the capabilities of your average CS student.
2. It is easy to write relative to C++. If you're in data science, those are your two options, so you choose Python.
Python is an extremely loose language - it's not restrictive at all in ways that matter. Is "you can't use two lines (unless you use a \)" restrictive when your language lets you arbitrarily replace functions at runtime?
It's not at all a readable language imo. It is very much write-optimized. It has tons of code golfing and one liners.
Compared to what? Python's closest peers are arguably the members of the P*-family of programming languages (Perl, Python, PHP, with Ruby as honorary member), and among that group, Python traditionally skews towards the readability end of the spectrum.
Agreed. This might have been applied to the "Zen of Python" days back in 1999, when the competition was Perl, but python has slowly morphed into a more and more perl-like language.
Like what? Where else does Python arbitrarily limit how many tokens may be part of a particular construct, or do anything else remotely similar?
I once believed that, but I now think it’s a terrible mistake. Python ends up encouraging classes instead but classes are a much worse abstraction than an anonymous function.