What I don't get is: why has Python 3 adoption been so slow? Is it just backward compatibility, or are there deeper problems with it that I'm not aware of?
What I don't get is: why has Python 3 adoption been so slow? Is it just backward compatibility, or are there deeper problems with it that I'm not aware of?
I can tell you about our situation.
We are an animation studio with decades of legacy Python 2 code. We sponser pycon and are one of the poster children for python.
We have absolutely no plans to switch to Python3.
Here are the various reasons:
- Performance is a big deal, and moving to a version of python that is slower is a no-go off the bat.
- Python3 has no compelling features that matter to us. The GIL was the one thing that should have been tackled in Python3.
- Since the GIL is here to stay, our long-term plan will likely involve removing more python from the pipeline rather than putting a huge effort into a python3 port.
- We have dependencies on 3rd-party applications (Houdini, Maya, Nuke) that do not support Python3
- We have no desire to port code "just because". Each production has the choice of either spending effort on Real Features that get pixels on the screen, or on porting code for No Observable Benefit. Real Features always win.
- Python3 has a Windows-centric "everything is unicode" view of the world that we do not care about. In our use case, the original behavior where "everything is a byte" is closer to UNIX. A lot of the motivation behind Python3 was to fix its Windows implementation. We are a Linux house, and we do not care about Windows.
- Armin's discussion about unicode in Python3 hits many of the points spot on.
Why has adoption been so slow? Simply because we have no desire to adopt the new version whatsoever. We'll be using Python2.7 for _at least_ the next 5 years, if not more.
It's far more likely that we'll adopt Lua as a scripting language before adopting Python3.
- We have dependencies on 3rd-party applications (Houdini, Maya, Nuke) that do not support Python3
This is our reason. It's a bit of a conundrum meets catch-22 situation. On one hand we would start using python3 if there would be a support for it, on the other hand no one wants to because of porting legacy code seems bothersome if everything works as it should. To be honest though, all of our python code is to augment those 3rd party applications. Everything we have that's not tied to those applications is C(99 more or less).
And you're welcome to do that.
But "we want a language frozen in time forever so we never have to maintain code" -- which seems to be what you're aiming for -- is not a goal you can achieve short of developing your own in-house language and never letting it make contact with the public (since as soon as it goes public it will change).
Meanwhile, the libraries are moving on and sooner or later they'll either move to Python 3, or be replaced by equivalent libraries with active maintenance, and the distros are winding down their support for Python 2. Switching to something else, and probably just rolling an in-house language you can control forever, is likely your only option if this is your genuine technical position.
The real feature needed is to eliminate the GIL. That would be worth breaking compatibility over.
Python 3.0 was slower than 2.7 due to several key bits being implemented in pure Python in 3.0. Since the 3.0 release (remember, Python's on 3.5 now) things that needed it have been rewritten in C, and as of Python 3.3 the speed difference is one. Also, on Python 3.3+ strings use anywhere from one-half to one-fourth the memory they used to.
As for "no compelling features", well...
* New, better-organized standard library modules for quite a few things including networking
* Extended iterable unpacking
* concurrent.futures
* Improved generators and coroutines with 'yield from'
* asyncio and async/await support in the language itself
* The matrix-multiplication operator supported at the language level (kinda important for all the math/science stacks using Python)
* Exception chaining and 'raise from'
* The simplified, Python-accessible rewritten import system
etc., etc., etc.
Given how many people and projects suddenly said "yeah, actually, we want to be on Python 3 now" after seeing the new stuff in 3.4 and 3.5, I think you're overestimating the number of people who don't care about these features.
1. Python 3.4 was released 7 years after 3.0. That's a very long time.
2. Assuming you're referring to the async features (none of the other stuff is really momentous), I can understand that for framework devs. It's really not a big deal for most users though, and because of the GIL, it's not as if they'll suddenly reap the benefits of parallelism. All it really means is nicer algorithm expressions and event loops.
3. There's no technical reason the new stuff couldn't have been added to Python 2. The roadblocks are manpower and politics.
4. Even if we stipulate async/await/asyncio are huge, busted Unicode support, 7 years of development, and breaking compatibility with everything is just a terrible tradeoff.
There's just no way this was a good idea, and Python devs could earn a lot of credibility back if they just said "oops". But there's no chance of that.
I wonder:
1. Would this have changed if there was a Python 2.8 with some new features (but still a GIL) that ran most 2.7 code?
2. Is your experience representative, i.e. are teams just not starting new projects in any version of Python, even though so many did in the last decade?
I am curious - if there was a GIL-removed version of Python 2, would you change your assessment that your "long-term plan will likely involve removing more python from the pipeline"? i.e. is the GIL the primary (or even sole) factor in that?
Perhaps you want to parse an old obscure file format, and the only code you find for it is from a usenet post in 1996. That code isn't being updated, and no one has ported it. That means you need to do the work to update it, and that can be hard when you aren't familiar with what it's supposed to do.
The other place I've seen people sticking with Py2 is when they've got a huge chunk of internal code. Some companies have been writing python for 20 years, and the original authors have long since left the company. It can be hard to write a business case for having someone spend several weeks updating all the old code, particularly if it's purely internal, and doesn't touch the internet.
The list I've gone off of is https://python3wos.appspot.com/. I use python a lot, and only now in late 2015 I might finally use python3 if starting a new project. When I started a new project last year, we used python2.
The biggest hold-out for me was gevent, which was only released 5 months ago. Gunicorn with gevent workers is my preferred stack for running python apps.
If you use protobufs or thrift, those both aren't yet on python3.
The wall of shame currently lists requests as not working on python3, though I think that might be a fluke.
These ones might not be a deal-breaker since you can have a separate environment for infrastructure & app code, but for some reason a lot of the infrastructure tools still haven't updated to python 3 (supervisor, ansible, fabric, graphite).
All together, it adds up to a not-insignificant number of things that aren't yet on python3. And even if nothing you use when you first start a project is python2 only, you have no idea what libraries you might need or want in the future and if those might be python2 only.
If you're willing to do a bit more than "pip install x", well I guess it doesn't work, back to py2, you can use almost everything on py3. (and yes requests is ported too)
Gevent said it looks like they're supporting Python3 as of the 1.1 release - http://www.gevent.org/whatsnew_1_1.html
I've used requests with Python3 quite a bit.. I have no idea why it's not listed as supported on the WoS, but their page shows it as working since 2012. https://pypi.python.org/pypi/requests/
Protobugs looks like it's now Py3 compatible, per the devs- https://github.com/google/protobuf/issues/646
ThriftPy has supported Python3 for a while (although ThriftPy is slower than the official lib). Apache Thrift has recently begun working with Python3 as well, however (https://issues.apache.org/jira/browse/THRIFT-1857)
I run Ansible/etc in their own virtualenvs, so they aren't part of system python for me, but you're right - I did have to write an Ansible module recently, and I recall that I did have to use Py2.
The only major package I recall having a problem with was PIL - Eventually I moved to Pillow as a drop-in replacement.
[1]: http://docs.python-requests.org/en/latest/
[2]: http://docs.python-requests.org/en/latest/#feature-support
Moreover, in machine learning Python 2.7 is still considered the default version of Python. (E.g.: some part of OpenCV need Python 2.7; until recently Spark supported only Python 2.7.)
Really this topic is coming to an end. Most libraries support Python 3 and if they don't there's better alternatives.
For new users and new projects there's no reason now except personal preference to choose Python 2 and in fact beginners who start with Python 2 are just instantly incurring a learning debt upon themselves to be paid down the track when they have to move to Python 3.
The community has some extremely vocal Python2 diehards but their arguments no longer hold water.
>>why has Python 3 adoption been so slow? Whatever, it's just history now.
import pymysql as MySQLdb
and everything just worked.> I'm also a fan of map and filter returning generators rather than being hard-coded to a list implementation
I find myself often having to wrap expressions in list(...) to force the lists. (which is annoying)
Generators make things much more complicated. They are basically a way to make (interacting, by means of side-effects) coroutines, which are difficult to control. In most use cases (scripting) lists are much easier to use (no interleaving of side effects) and there is plenty of memory available to force them.
Generators also go against the "explicit is better than implicit" mantra. It's hard to tell them apart from lists. And often it's just not clear if the code at hand works for lists, or generators, or both.
So IMHO generators by default is a bad choice.
> stream fusion is a good thing
I don't think generators qualify for "stream fusion". I think stream fusion is a notion from compiled languages like Haskell where multiple individual per-item actions can be compiled and optimized to a combined action. Python instead, I guess, just interleaves actions, which might even be less efficient for complicated actions.
Out of curiosity, why do you need to force the lists?
> Generators make things much more complicated. They are basically a way to make (interacting, by means of side-effects) coroutines
Huh? Generators are a way to not make expensive computations until you have to--as well as to not use memory that you don't need. Basically, if all you're doing with a collection of items is iterating over it (which covers a lot of use cases--but perhaps not yours), you should use a generator, not a list--your code will run faster and use less memory.
> In most use cases (scripting) lists are much easier to use (no interleaving of side effects) and there is plenty of memory available to force them.
Generators don't have to have side effects. And there are plenty of use cases for which you do not have "plenty of memory available" (again, perhaps not yours).
> IMHO generators by default is a bad choice.
I think lists by default was a bad choice, because it forces everyone to incur the memory and performance overhead of constructing a list whether they need to or not. The default should be the leaner of the two alternatives; people who need or prefer the extra overhead can then get it by using list() (or a list comprehension instead of a generator expression, which is just a matter of typing brackets instead of parentheses).
for i in range(large_number):
...
(similarly enumerate, zip etc)i.e., where you don't bind the generator to a variable, but instead immediately consume and reading from it as an iterator has no other side-effects.
And that's the only usage of generators that in my usage is both common and practical. Other use has always quickly become a mess for the above named reasons.
I don't disagree generators can be occasionally useful (or often, for your special applications). But mostly it's a pain that they are the only thing many APIs return and that they are not visually distinctive (in usage) from lists and iterators. For example
c = sqlite3.connect(path)
rows = c.execute('select blablabla')
do_some_calculation(rows)
print_table(rows)
Gotcha! "rows" was probably already empty when print_table was supposed to print it.
But how can you know? Hunt down all the code, see what the functions do and what they want to receive (lists, iterators, generators? probably they don't even know). And what if the functions change later? Even subtler bugs occur if the input is consumed only partly.So by far the common (= no billions of rows) sane thing to do is
rows = list(c.execute('select blablabla'))
Which is arguably annoying and requires a wrapper for non-trivial things.Or, to put it another way, since there are two possibilities--realize the list or don't--one of the two is going to have to have a more verbose spelling. The obvious general rule in such cases is that the possibility with less overhead is the one that gets the shortest spelling, i.e., the default.
Also, if you've already realized a list, it's too late to go back and un-realize it, so there can't be any function like make_generator(list) that saves the overhead of a list when you don't need it. So there's no way to make the list alternative have the shorter spelling and still make the generator alternative possible at all.
As far as your sqlite3 example is concerned, why can't do_some_calculation(rows) be a generator itself? Then print_some_calculation would just take do_some_calculation(rows) as its argument. Does the whole list really have to be realized in order to print the table? Why can't you just print one row at a time as it's generated?
Basically, the only time you need to realize a list is if you need to do repeated operations that require multiple rows all at the same time. But such cases are, at least in my experience, rare. Most of the time you just need one row at a time, and for that common case, generators are better than lists--they run faster and use less memory.
If you need to do repeated operations on each row, you just chain the generators that do each one (similar to a shell pipeline in Unix), as in the example above. This also makes your code simpler, since each generator is just focused on the one operation it computes; you don't have to keep track of which row you're working on or how many there are or what stage of the operation you're at, the chaining of the generators automatically does all that bookkeeping for you.
'%s %s' % ('one', 'two')
With '%s %s'.format('one', 'two')
The latter is just more annoying to type. Stupid argument I know but I find myself grumbling to myself every time... '{} {}'.format('one', 'two')
Either way, the former example still works in Python 3.5, so that syntax hasn't gone away. `format` is preferred, though. This Stack Overflow question has some good answers as to why: http://stackoverflow.com/questions/5082452/python-string-for... '{} {}'.format('one', 'two')
Your experience may vary, but in my experience, when I switched to .format(), I found a number of bugs in code that used % instead. As mentioned, you can continue using %.I love love love printf() and it's ilk, so switching to something else (no matter how well designed) seemed asinine at first.
But .format() has really started to grow on me and did uncover some subtle bugs in old code.
'Coordinates: {latitude}, {longitude}'.format(latitude='37.24N', longitude='-115.81W') f'{one} {two}'https://docs.python.org/2/library/string.html#string.Templat...
Another reason is that the advantage of switching is just not that big, if you already have everything working in Python2.
I don't think the lesson is "never break compatibility". The lesson is "don't compete with yourself by releasing a product that is actively worse than your current version"
Within the last six months we've moved to writing all new code in Python 3 and migrating a fair bit of legacy code as well. Been fairly smooth on Linux -- a bit rockier on Windows.
This wasn't just down to the fact that Python 3 didn't support u"" (it was 3.3 that was added, for reference), but also down to the fact that much of the eco-system still supported RHEL5's default of Python 2.4 which meant `from __future__ import unicode_literals` wasn't an option (it is, almost certainly, a less good option, but it's in many ways good enough).
It took a few versions of 3 to hit a sweet spot (in some cases features that were removed in the initial version of 3 have slowly been re-added in subsequent versions). There were a lot of crucial libs that needed to be ported. Just general inertia.
Following their glacial release schedule, maybe we'll see Python 3 by 2019 in RHEL 8.
(It of course limits software that is to be shipped with the distro itself)