PEP 492 – Coroutines with async and await syntax
python.org
python.org
The greatness of Python would have been impossible without rejecting virtually all feature requests. To move the vision of Python forward, it was necessary to depart with old syntax.
Python 3 is happening. Deal with it.
Moving a large codebase to v3 is expensive for almost no benefits? Some parts of the community already moves on to Go, etc. Many third party libraries are v2 only and unmaintained. Many languages that broke backwards compatibility ultimately failed (their community moved on) or split their community. Examples: Modula 2/3/Oberon4, Perl 4/5/6, Lua 5.1(LuaJIT)/5.2, Visual Basic 6/.Net, QBasic/Visual Basic 1, VBA/VBA.Net, J++/J.Net/C#, etc.
Fibers and coroutines are key features of certain language and API's for decades. WinAPI16 had already Fibers (MS Word), Lua's coroutines, Go's goroutines, etc.
Also Oberon had a different purpose than Modula-2.
In Oberon's case, you have actually Oberon, Oberon-2, Component Pascal, Active Oberon, Zonnon and Oberon-07.
It is not the same as the next version of a given language, rather different approaches how to design memory safe languages for systems programming.
One does not see such amount of discussions regarding other languages that suffer even bigger transitions problems.
Java is now version 8, still there are lots of places having to deal with versions 4 and 5.
.NET is getting 4.6/5 version in the upcoming months, still many enterprises still have version 3.5 code bases.
Everyone is discussing the benefits of C++14, while many corporations still use pre-C++98 like code.
Yet in the Python community it is such a big deal.
http://www.oracle.com/technetwork/java/javase/8-compatibilit...
Breaking changes in Java 7
http://www.oracle.com/technetwork/java/javase/compatibility-...
Breaking changes in Java 6
http://www.oracle.com/technetwork/java/javase/compatibility-...
Breaking changes in Java 5
http://www.oracle.com/technetwork/java/javase/compatibility-...
Breaking changes in all Java versions up to 1.4
http://www.oracle.com/technetwork/java/javase/compatibility-...
I no longer remember which version it was and don't feel like going through those lists now, but I remember one of those versions changed some JDBC interfaces which would break any application using JDBC.
Python had its share of breaking changes as well over the years and there wasn't much fuss about them. Who refused to upgrade over class name(object) or say hex(-1) producing '-0x1' instead of '0xffffffff'?
lol wut? How exactly does adding a method to a JDBC interface "break any application using JDBC"?
Maybe I should have written extending the JDBC classes instead.
Just like nobody walks around and complains that new versions of the servlet API break any Java web application because web applications are not supposed to implement the servlet API. That's the job of the server.
Another use case, many developers mock JDBC by creating their own dummy drivers, specially in large companies where mocking libraries are frown upon.
Nope, they are database specific.
So we went from "would break any application using JDBC" to large company rolling their own database drivers for reasons that can't be disclosed (probably to protect the guilty).
> as it was the case in the application I had to fix. Why it was made so, will be kept under the covers of the NDA agreement.
Standard case of a large company doing it wrong and blaming somebody else when it comes back to bite them rather than fixing it. And hiding behind an NDA. Seriously what was the expectation? That JDBC never changes? Because at that time JDBC was already at version 3.0 which is obviously the final version after which no features would ever be added again.
> Another use case, many developers mock JDBC by creating their own dummy drivers, specially in large companies where mocking libraries are frown upon.
If you're doing it wrong then you're doing it wrong. Nobody to blame but yourself. Just because you're a large company doesn't make it right. Part of why it's wrong is that it will come back and bite you later on. And that's the problem in this case? The compiler tells you where you need to implement which methods.
Two give you two other examples. Interfaces for HTTP likely need to be updated when something in HTTP changes (WebSockets, HTTP/2). Servers implementing this will need to be updated to implement this. You don't accuse the language of the web server to make a breaking change. That's just how these things work. Same for SSL/TLS features like SNI. Sure it could be that you absolutely have to run your own web server or SSL/TLS implementations for reasons that you can't disclose because you're under an NDA. But then you expect that you'll have to maintain them and add additional features, don't you?
Ruby 2.0 had similar levels of incompatibility; I think people simply didn't have those large enterprise codebases in Ruby to make a fuss about.
Java, .net and C++ go to extreme efforts for backwards compatibility, compromising their current/future versions as a result.
True, but they still introduce breaking changes as well.
I just posted all the Java release notes in a sibling post, and could do the same for .NET and C++.
Maybe the amount of breakage isn't as big as Python 2 -> 3, but it does happen.
Shame on the community for not learning the lessons of Perl. Fragmentation is far worse than an imperfect language.
I know this is just the situation in my corner, and that it looks rosier for other users - e.g. for web development, system administration, etc..
The thing is, we have decades worth of legacy code - large libraries and small config scripts - that no one is going to rewrite. All python versions up till 2.7 were mostly backwards-compatible, and we came to depend on that. That the changes towards 3.x were breaking seems totally unneccessary (except for the string/unicode thing), which also gives it a psychological component in my opinion. Give us back the stupid print statement, and I bet this alone will massively increase adoption.
And if someone makes a patch to Python to import a Python 2 module in Python 3 (old syntax, old str/unicode, and a copy of the old stdlib - and it doesn't matter if it is 10% slower or never gets merged upstream), I'd be willing to pay $$ for it.
You might be able to require parens for methods with no arguments, but optional parens with arguments. It might require some clever hacks in the parser, but I suspect it's possible. The bigger question is whether or not the result would be Pythonic.
len flurp
to quickly get the length of an array.For the people that are stuck with Python 2 legacy code, we'd need something else.
Python-future [1] seems to be able to import some Python 2 modules, but I'm not sure how far it goes.
>>> print "hello"
SyntaxError: Missing parentheses in call to 'print'
If it can print a SyntaxError explaining the problem then it can print "hello".The only snag is that a tuple would be printed slightly differently, but that is harmless.
People should be upset that a new release of Python breaks their code, because the Python developers acted unprofessionally in expecting all python code in existence to be rewritten to suit their aesthetic fetishes.
You're really foolish if you believe that's better. The print keyword isn't what is preventing 3k adoption.
https://pypi.python.org/pypi/python-bond
> You can freely mix Python versions between hosts/interpreters (that is: you can run Python 3 code from a Python 2 host and vice-versa).
But more importantly, keep in mind that many academics move on to a new institution every couple of years (after finishing undergrad/Ph.D./post-doc, etc.), so any code written more than 5 years ago likely has no maintainer and no one in the lab knows how it works.
My dad is a chemist, and I recently learnt that some Fortran code he wrote in the nineties is till being used in academia.
Try to explain to his colleagues that you need an on-site engineer to maintain python code just a decade old.
Anyhow, while I disagree with the need of a dedicated engineer to port code to python3, porting isn't much of a big deal. If your dependencies work (eg: your libraries are python3-compatible), most of you work is done by 2to3. Very little effort is needed after that.
The problem up to now, has been waiting for you dependencies/libraries to achieve python3-compatibility (recursively, of course). But we've already moved past that.
I highly doubt that "End of life" means you can't get it to run anymore.
Very true, but this happens regardless of whether it was Python2.x, Python3 or MATLAB code.
The problem is that many academics writing scientific code rarely have the training/education/experience to write maintainable code (if you feel this does not apply to you, you are probably the exception, and if you ever shared code with colleagues, you are aware how rare your skill is in academia). Also most of it is or started out as, quick experiments, try-outs, for that, and some other (even political) reasons, there is not a lot of incentive to write beautiful maintainable code.
But then again, a lack of reproducibility sometimes seems to be what it takes to succeed in Academia...
my standing question in the debate has always been what is the big deal with having built a robust 2to3 preprocessor?
i figure if you make a change to the spec migrating existing code to the new syntax should be a pretty straight forward ifttt, something that should be a necessary addition to the change, a la tests
knowing full well 'pretty straight forward' is the bane of all software endeavors could anyone explain to me what is going on within the scope of 2to3 preprocessors?
and why such preprocessors seem an impossible boon very few, except a handful of seeming independants, endeavor
y = yield f(x)
z = yield g(y)
w = yield h(z)
...
In this case it would be async/await instead of pure yields.The worst thing was having to hunt for Twisted version of libraries. "Oh you want to talk XMPP? Nah, can't use this Python library, have to find the Twisted version of it". It basically split the library ecosystem. Now presumably it will be having to look for async/await version of libraries that do IO.
f(x, function(y) {
g(y, function(z) {
h(z, function(w) {
do_something_with_w(w);
});
});
});
What would be a better syntax? The Java/C++ way of threads & hidden shared-state concurrency is a total mess; it hides all the potential yield points, so you never know when your flow of control might block or what shared state might need locking. Channels in Golang are better - they at least have some syntactic support - but the reification of the yield point into a concrete channel can end up creating a fair bit of boilerplate in the common case where you make a bunch of one-off async RPCs to remote services. Maybe Erlang has it right, where the pid is an implicit channel you can send messages to - but then you still need to pass that into any async library function, and it gives you no typing discipline. Maybe we really need something like futures where the syntax: y <- f(x)
z <- g(y)
w <- h(z)
means "wait for the promise returned by expression f(z) to resolve, returning control to the executor, and then assign the result to y."The splitting-the-library-ecosystem thing is a big part of it too (and why ES6 won't magically fix the Node ecosystem), but that's why Python is putting async into the language & standard library itself. At least then the stdlib will support it, and there will be strong social pressure to use the same concurrency mechanism in all libraries.
> z <- g(y)
> w <- h(z)
So you would basically want "<-" instead of "yield from"? Or is there something additional I'm missing? My main thing against this, is that it's very opposed with general Python principles. Think "||" vs "or", "&&" vs "and", and "test ? value : alternate" vs "value if test else alternate".
It is definitely a bit verbose, but I decided that the clarity for the rest of the code is worth putting a yield before each function call. Also I've found a few projects (for Tornado at least) that cut down on this boiler plate and make the yield only required at the lowest level where the async really happens. [0]
> The worst thing was having to hunt for Twisted version of libraries. "Oh you want to talk XMPP? Nah, can't use this Python library, have to find the Twisted version of it". It basically split the library ecosystem. Now presumably it will be having to look for async/await version of libraries that do IO.
I work with Tornado and this is absolutely the worst part. At least with Tornado the newest version is embracing interoperability with python 3.4+ native AsynIO.
I don't believe coroutines/async routines will ever be practical without decorators, so you might as well keep them and not disturb one of the most basic syntax rules of Python.
Why do you believe so?
Have you seen the implementation of asyncio.coroutine? Have you read the PEP thoroughly and saw the downsides of using a decorator/generators?
edit: I can't explain topics better than I did so in the rationale section of the PEP. If I could, I would have explained it better in that section in the first place ;)
As for asyncio.coroutine decorator -- it's just a very simple wrapper, that makes sure that the decorated object is a generator-function. If it's not -- it wraps it in one.
It also does some magic to enable debug features. But with some serious shortcomings (that is also explained in the PEP in great detail).
My point is: there is absolutely no other value in that decorator. There is nothing fundamental that it does, it just fixes the warts. Documentation, tooling, "easier to spot", etc arguments are unfortunately weak.
The PEP makes coroutines in python a first-class language concept, with all the benefits you can have from it (better support in IDEs, tooling, sphinx, less questions on StackOverflow).
Disclaimer: I'm the PEP author. I'm also python core developer, and I contributed to asyncio a lot.
Most of this is just "feeling". But good Python makes me feel good, so I sort of expect that from new features if I am to vote them up.
I might like "codef" more than "async def", if there absolutely have to be first-class coroutines.
I thereby ask: is there something about Python 3 that makes this feature extremely easier to implement or extremely easier to integrate? If not, wouldn't it be more interesting to have "more impact" by improving the lives of a larger number of people? Python 3 was maybe an interesting experiment, but given how well Ruby 1.8->1.9 went and how poorly Perl 5->6 has been, it seems like people should be learning some lessons and fighting a more winnable battle.
Maybe the people who want to encourage people to migrate to Python 3 should he spending their time not providing "killer features" but instead narrowing the gap between the two versions, not by back porting features to Python 2.7, but by making it easier to use code written for Python 2 in Python 3. It should not have taken until 3.3 to see the u'' syntax return for compatibility with Python 2. :(
The current strategy of "try to convince an army of people to port a bunch of code from Python 2 to Python 3 while people are at the same time still often writing code for Python 2", which is what we are seeing being asked of people in other posts here over the last few days (there was a post about this for Debian) just seems like a waste of effort leading to tons of lost ground to alternative languages.
Python 3 is not Perl 6, and that's a banal comparison to make. Python 3 has had many releases, and has had most packages ported to it. Perl 6 has had no releases. Ruby 1.9 should be an example that making incompatible changes that force everyone to upgrade is possible as long as you make timely releases.
What gap do you still think there is to be narrowed? We're not on 3.2 anymore. Complaining that "it should not have taken until 3.3" doesn't matter now.
What most people are arguing is quite the opposite of you -- they're saying they could switch to Python 3, and they would if there were a killer feature that mattered to them, but just being cleaner and making the core developers happier isn't enough for them to change the way they use Python. So I say, bring on the killer features.
Now, I know that open-source projects are sometimes quite conservative with versioning, and one project's 0.9 may be more stable than another's 2.0. (I maintain an HTML parser [1] that's still on 0.9.3 and yet is more robust and better tested than one that is on 3.8.2.) Is this actually the case with Rakudo, though? You can write real production software with Python 3.4; can you with Rakudo #86?
https://mail.python.org/pipermail/python-list/2008-December/...
Python 2.x is not going anywhere soon. The idea that is going to become unmaintained is ridiculous. If you like it, keep using it, I say. However, maybe take a look at the "What's new in Python 3.x" documents sometime. There are a lot of nifty new features, even if none of them are "killer".
> I don't want to switch to Python 3 because of the changes to the language. You should make this change to Python 2.
You can't have it both ways. If you prefer Python 2 because the language hasn't changed, then don't demand changes to Python 2, and certainly don't accuse people of "fighting ideological battles" just because they are adding new features to the latest version of the language instead of old versions.
Actually, no. Guido mentioned in PyCon 2015 keynote from last week that only ~10% of packages from pypy have been ported. If I remember correctly there's ~50,000 packages, and only ~5,000 have been ported.
The fact that numpy, scipy and pandas are on green means a lot. They used to be a massive blocker. The secondary stops I've seen first hand were matplotlib and psycopg2.
A curious side-effect of this PEP might be that lack of py3-gevent could become less of an issue. At least for fresh and new projects that require co-operative concurrency. From a purely personal point of view - Oauth2, protobuf and newrelic plugin-agent are probably the biggest gaps.
Overall the list looks pretty good. I'm curious to hear what packages other heavy python users consider as blockers for even considering to migrate to python 3.
The gap has been getting smaller. Python 3.5 will include %-style formatting for byte strings. I helped implement that feature since I feel it will make porting code easier (plus it makes some programming tasks, like network protocols, easier). The strict handling of bytes/unicode is what really trips up people from porting and there is just no good way to further smooth that path, IMHO.
I feel the upgrade path has been handled badly. People were wildly optimistic about how fast people could port code and much more effort should have been spent on making the transition easier. A little too much purity instead of practicality. For example, u'foo' style strings were not accepted until Python 3.3. That just made porting more difficult for great reason.
Given the massive amount of Python 2.x code in the world, I fully expect that version to live much longer than most people expect. There is still old Fortran and Cobol code out there for example. Businesses don't want to rewrite a working system. Someone is going to keep making releases of the 2.x branch. Maybe it will be a fork.
At this point, I'd bet money on Python 3 succeeding. The vast majority of users still use 2.7 but there is steady progress of porting. The 3.x branch has received a lot of new and useful improvements and that is where all the development is happening. The memory savings from the new string representation will help my projects a lot. As I said elsewhere, asyncio is really sweet. The core team needs to keep the improvements coming. More carrots and less sticks, IMHO.
A lot of new python programmers learn python3 first, and are extremely hostile to people using python2.
It doesn't help at all. If you love python3, the best thing you can do is help make it more awesome!
Shouting at people using python2 or telling them to 'deal with it' makes you look like a jerk. Don't be a jerk, it's bad karma.
With many important code bases such as Ansible and Django you still need to keep a Python 2 around. (Yes, I know Django works on 3, but less people use it and performance is slightly lower.) So you are dual stacked for the forseeable future if you choose the Python 3 route.
So the Python ecosystem is a little more complicated than "old" Python 2 code versus "new" Python 3 code. Should I put my consultant hat on I would still advise people to start new developments with Python 2, but to keep an eye out for future compatibility issues. That recommendation hasn't changed in five years.
Python 2 will be kept around until 2020. Obviously there will still be companies using it after that date. But they have been given ample notice.
I use Django with Python 3 and it's great. Admittedly, I don't need a lot of performance though.
I really like Python 3 mainly for its more explicit treatment of Unicode.
I believe the only way for Python to remain a success is if the community adopts Python 3 and puts in the effort to migrate away from 2. Once the community is there, organizations will follow.
I'm all for this change, but if people were willing to look at asyncio, they'd see this is largely a syntactic tweak to that.
a, b, c = atools.join([f(), g(), h()])
and select among multiple async tasks (e.g. to implement a timeout): with atools.select({"res": f(), "timeout": timeout(10)}) as tasks:
res, which = await tasksI like the idea of 'a, b = await foo(), bar()' syntax. I'll think of how we can implement it and integrate to the PEP.
Thanks!
it is easy to confuse coroutines with regular generators, since they share
the same syntax; async libraries often attempt to alleviate this by using
decorators (e.g. @asyncio.coroutine [1] );
it is not possible to natively define a coroutine which has no yield or
yield from statements, again requiring the use of decorators to fix
potential refactoring issues;
support for asynchronous calls is limited to expressions where yield is
allowed syntactically, limiting the usefulness of syntactic features, such
as with and for statements.
[1] https://www.python.org/dev/peps/pep-0492/#rationale-and-goal...I just can't somebody somebody would like to introduce a whole new syntaxe just for dislike of decorator and 2 edge cases twisted/tornado/greenlet lived with for years. So I though I missed something. Espacially since we have ways to do async and "there should be one way to do it" is as much important as keeping builtins number down.
So I'm asking again : am I missing something or is it again on of these "I don't like it so let's do it like in language x" claims like we had for brackets, lambdas et alike ?
Victor Stinner: """As a contributor to asyncio, I'm a strong supporter of this PEP. IMO it should land into Python 3.5 to confirm that Python promotes asynchronous programming (and is the best language for async programming? :-)) and accelerate the adoption of asyncio "everywhere". ¶The first time Yury shared with me his idea of new keywords (at Pycon Montreal 2014), I understood that it was just syntax sugar. In fact, it adds also "async for" and "async with" which adds new features. They are almost required to make asyncio usage easier. "async for" would help database ORMs or the aiofiles project (run file I/O in threads). "async with" helps also ORMs."""
Marc-Andre Lemberg: """Thanks for proposing this syntax. With async/await added to Python I think I would actually start writing native async code instead of relying on gevent to do all the heavy lifting for me ;-)"""
Guido: """I'm in favor of Yuri's PEP too, but I don't want to push too hard for it as I haven't had the time to consider all the consequences."""
etc.
For more: https://mail.python.org/pipermail/python-ideas/2015-April/th...
A Python generator is a lazily computed, readable stream. Generators have been integrated very nicely with the rest of the language (the iterator protocol, for loops, etc.). The result is that Python now has excellent support for the style of programming that builds a system out of composable components.
It is almost incidental that a generator function is really a kind of coroutine. This fact does make them very convenient to write; but when generators were designed, I don't think anyone was saying, "Let's add a high-quality coroutine facility to Python."
But somewhere along the way, someone noticed that generators are coroutines. If you don't mind doing a "yield" whenever a coroutine needs to give up control -- possibly using a bit of twisted syntax to make that work -- generators are in fact coroutines in full generality.
But they don't look that way. If you just want to pass around control and do things asynchronously, you are still required to use the send-this-value-out-on-the-stream syntax. Conceptually, this is not the right abstraction. It is also occasionally cumbersome; wrapping a code snippet in some asynchronous processing is messy, for example.
That, plus the fact that the async/await (or future/promise) idea has proven itself to be effective and useful in other languages, suggests that Python needs additional constructions for doing coroutines.
That being said, I'm a bit leery of this PEP myself. The Rationale and Goals section, posted by ceronman, is all about deficiencies in the syntax. The real problem is that Python syntax is not sufficiently versatile to allow a programmer to create needed abstractions. We're talking about adding two new keywords and several constructions to support abstractions whose functionality can be implemented currently in Python, but cannot be given a nice interface.
I would prefer it if attention were given to the kind of syntax Python allows for programmer-created abstractions. Give me enough power that I can write async & await myself, along with some other things I might come up with, but which no one else has thought of yet. Polished versions of async/await can then be added to the standard library.
If you are using asyncio (or something similar), you will always need to pay attention to the "concurrent choreography" (for lack of a better word). I find I rarely confuse coroutines with other routines, and it's mostly when calling asyncio based third party libraries.
I'm curious, is this just my warped perception? Did this terminology and feature set previously exist in programming language theory or practice? I know F# had "async" for a while before C#. Was there "prior art" that looked basically the same in Lisp, Haskell, Scala, ML, etc. or in some research language long before that?
A bunch of modern C# features are actually lifted from F#, which had them first on the .NET platform. Sometimes the C# version is simplified or has some of the sharp edges filed off (which makes it less powerful, but safer). Async/await is definitely one of these.
"Confessions of a used programming language salesman (getting the masses hooked on haskell"
http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.118....
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/n445...
a. this is functionally equivalent to coroutines implemented via library support through generators
b. Python will never ever get rid of the GIL, i.e. won't have real parallelism (given that the JVM/CLR Python implementations are basically dead and PyPy is chasing the pype dream of STM, we're left with CPython the only actual implementation)
Is there any reason to add new syntax? Other than wishful thinking that one day someone will write a Python runtime with proper JIT and parallelism support?
P.S - Ruby's issues system is hard to follow.
I'm not sure that special syntax is necessary for widespread adoption of Fibers, since Ruby syntax is flexible enough to provide a sort of DSL for working with Fibers.
Sadly, it died some years ago.
[0] https://www.ruby-lang.org/en/news/2003/12/19/new-ruby-change...
PEP 492, from what I understood, is more about streamlining the syntax, and the major differences between asyncio (pre this proposed syntax) and green threads in terms of programming model still apply.
I've done a project making heavy use of Python 3's asyncio feature. It is much nicer than doing the same job with threads or call-backs, in my experience. I'd welcome nicer syntax but the current functionality seems to work well.
I can't understand why someone would design a language in 2009 (Node.js) that uses callbacks. For simple programs it works but it doesn't take long to end up with an unmaintainable mess (in my experience).
* https://review.openstack.org/#/c/153298/
* http://techspot.zzzeek.org/2015/02/15/asynchronous-python-an...
* http://lists.openstack.org/pipermail/openstack-dev/2014-Febr...
And others...
Asyncio "threads", on the other hand, never switch context between "yield from" instructions. This makes it usable with the most bungled and complicated C APIs (e.g. Blender...), and also lets us reason more clearly about the concurrent behavior.
Also, this kind of message-passing concurrency is not meant for cpu-intensive work, but rather for situations where there's nothing to do in a particular task much of the time. Concurrent threads/tasks in frameworks like NodeJs, Tornado and Asyncio are very good at doing nothing.
For example in a web server, a traditionally threaded "worker" spends a lot of time waiting for a connection, waiting for database requests to finish, and so on. The operating system can use that time to execute other threads. But the memory of the thread sits idle during that time.
Scheduling in asyncio is not imposed by the operating system or the programmer, and only partly by the event loop, but largely because of external circumstances like network latency/bandwith, how long do the database queries take, and how fast the user is acting.
As a comparison to glyph http://techspot.zzzeek.org/2015/02/15/asynchronous-python-an... (the creator/author of sqlalchemy) is a pretty good read :-)
I have written quite a lot of asyncio and tornado code by now, and I can't remember a single instance where I had to yield/yield from/wait just because the computation would freeze up the thread or increase latency.
Coroutines really aren't about explicit scheduling, at the very most you can specify what code is executed without switching. And the latter is absolutely necessary in some situations. For example, threads or greenlets freak out Blender.
I love you.
I meant the above comment in earnest, not sarcastically. Python 3.* lets the people who like that sort of thing have their cake and eat it too (and I am one of the people who drools at new shiny tech.) But the fact that Python 2.7 is stable for the foreseeable future is something to celebrate in my opinion, not condemn.