Async Python is not faster
calpaterson.com
calpaterson.com
When you're dealing with external REST APIs that take multiple seconds to respond, then the async version is substantially "faster" because your process can get some other useful work done while it's waiting. Obviously the async framework introduces some overhead, but that bit of overhead is probably a lot less than the 3 billion cpu cycles you'll waste waiting 1000ms for an external service.
As I describe in the first line of my article I don't think that people who think async is faster have unreasonable expectations. It seems very intuitive to assume that greater concurrency would mean greater performance - at least one some measure.
> When you're dealing with external REST APIs that take multiple seconds to respond, then the async version is substantially "faster" because your process can get some other useful work done while it's waiting.
I'm afraid I also don't think you have this right conceptually. An async implementation that does multiple ("embarrassingly parallel") tasks in the same process - whether that is DB IO waiting or microservice IO waiting - is not necessarily a performance improvement over a sync version that just starts more workers and has the OS kernel scheduler organise things. In fact in practice an async version is normally lower throughput, higher latency and more fragile. This is really what I'm getting at when I say async is not faster.
Fundamentally, you do not waste "3 billion cpu cycles" waiting 1000ms for an external service. Making alternative use of the otherwise idle CPU is the purpose (and IMO the proper domain of) operating systems.
Sure, the operating system can find other things to do with the CPU cycles when a program is IO-locked, but that doesn't help the program that you're in the situation of currently trying to run.
> An async implementation that does multiple ("embarrassingly parallel") tasks in the same process - whether that is DB IO waiting or microservice IO waiting - is not necessarily a performance improvement over a sync version that just starts more workers and has the OS kernel scheduler organise things. In fact in practice an async version is normally lower throughput, higher latency and more fragile. This is really what I'm getting at when I say async is not faster.
You're right. "Arbitrary programs will run faster" is not the promise of Python async.
Python async does help a program work faster in the situation that phodge just described (waiting for web requests, or waiting for a slow hardware device), since the program can do other things while waiting for the locked IO (unlike a Python program that does not use async and could only proceed linearly through its instructions). That's the problem that Python asyncio purports to solve. It is still subject to the Global Interpreter Lock, meaning it's still bound to one thread. (Python's multiprocessing library is needed to overcome the GIL and break out a program into multiple threads, at the cost that cross-thread communication now becomes expensive).
This isn't how it works. While Python is blocked in I/O calls, it releases the GIL so other threads can proceed. (If the GIL were never released then I'm sure they wouldn't have put threading in the Python standard library.)
> Python's multiprocessing library is needed to overcome the GIL
This is technically true, in that if you are running up against the GIL then the only way to overcome it is to use multiprocessing. But blocking IO isn't one of those situations, so you can just use threads.
The comparison here is not async vs just doing one thing. It's async vs threads. I believe that's what the performance comparison in the article is about, and if threads were as broken as you say then obviously they wouldn't have performed better than asyncio.
--------
As an aside, many C-based extensions also release the GIL when performing CPU-bound computations e.g. numpy and scipy. So GIL doesn't even prevent you from using multithreading in CPU-heavy applications, so long as they are relatively large operations (e.g. a few calls to multiply huge matrices together would parallelise well, but many calls to multiply tiny matrices together would heavily contend the GIL).
> No it's not, just use threads.
I just wanted to expand on this a little to describe some of the downsides to threads in Python.
Multi-threaded logic can be (and often is) slower than single-threaded logic because threading introduces overhead of lock contention and context switching. David Beazley did a talk illustrating this in 2010:
https://www.youtube.com/watch?v=Obt-vMVdM8s
He also did a great talk about coroutines in 2015 where he explores threading and coroutines a bit more:
https://www.youtube.com/watch?v=MCs5OvhV9S4&t=525s
In workloads that are often "blocked" like network calls our I/O bound work loads, threads can provide similar benefits to coroutines but with overhead. Coroutines seek to provide the same benefit without as much overhead (no lock contention, fewer context switches by the kernel).
It's probably not the right guidelines for everyone but I generally use these when thinking about concurrency (and pseudo-concurrency) in Python:
- Coroutines where I can.
- Multi-processing where I need real concurrency.
- Never threads.
The point is, many people think (including you judging by your comment, and certainly including me up until now but now I'm just confused) that in Python asyncio is better than using multiple threads with blocking IO. The point of the article is to dispel that belief. There seems to be some debate about whether the article is really representative, and I'm very curious about that. But then the parent comment to mine took us on an unproductive detour that based on the misconception that Python threads don't work at all. Now your comment has brought up that original belief again, but you haven't referenced the article at all.
Threads and async are not mutually exclusive. If your system resources aren't heavily loaded, it doesn't matter, just choose the library you find most appropriate. But threads require more system overhead, and eventually adding more threads will reduce performance. So if it's critical to thoroughly maximize system resources, and your system cannot handle more threads, you need async (and threads).
Absolutely false. OS threads are orders of magnitude lighter than any Python coroutine implementation.
During any operation that does not need to modify Python objects, it is safe to unlock the GIL. Yielding control to the OS to wait on I/O is one such example, but doing heavy computation work in C (e.g. numpy) can be another.
Single-threaded performance is a major issue as that's most Python code.
Any solution which wants to consider per-object locking has to consider removing refcounting, or locking the refcount bits separately, as locking/unlocking objects to twiddle their refcounts is going to be ridiculously expensive.
Ultimately, the Python ownership and object model is not condusive to proper threading, as most objects are global state and can be mutated by any thread.
Workers (usually live in a new process) are not efficient. Processes are extremely expensive and subjectively harder for exception handling. Threads are lighter weight..and even better are async implementations that use a much more scalable FSM to handle this.
Offloading work to things not subjective to the GIL is the reason async Python got so much traction. It works really well.
On the web when the bulk of your application code time is waiting on APIs, database queries, external caches or disk I/O it creates a dramatic increase in the capacity of your server if you can do it with minimal RAM overhead.
It's one of the big reasons I've always wanted to see Techempower create a test version that continues to increase concurrency beyond 512 (as high as maybe 10k). I think it would be interesting.
Python doesn't block on I/O.
If you go back to the origins of Erlang, the intent was to build a language that would make it easier to write software for telecom (voice) switches; what comes out of that is one process for each line, waiting for someone to pick up the line and dial or for an incoming call to make the line ring (and then connecting the call if the line is answered). Having this run as an isolated process allows for better system stability --- if someone crashes the process attached to their line, the switch doesn't lose any of the state for the other lines.
It turns out that a 1980s design for operational excellence works really well for (some) applications today. Because the processes are isolated, it's not very tricky to run them in parallel. If you've got a lot of concurrent event streams (like users connected via XMPP or HTTP), assigning each a process makes it easy to write programs for them, and because Erlang processes are significantly lighter weight than OS processes or threads, you can have millions of connections to a machine, each with its own process.
You can absolutely manage millions of connections in other languages, but I think Erlang's approach to concurrency makes it simpler to write programs to address that case.
Immutable data, heap isolated by concurrent process and lack of shared state, combined with supervision trees made possible because of extremely low overhead concurrency, and preemptive scheduling to prevent any one process from taking over the CPU...create that operational consistency.
It's a combination of factors that have gone into the language design that make it all possible though. Very big and interesting topic.
But it does create a significant capacity increase. Here's a simple example with websockets.
https://dockyard.com/blog/2016/08/09/phoenix-channels-vs-rai...
This is very true, especially when actual work is involved.
Remember, the kernel uses the exact same mechanism to have a process wait on a synchronous read/write, as it does for a processes issuing epoll_wait. Furthermore, isolating tasks into their own processes (or, sigh, threads), allows the kernel scheduler to make much better decisions, such as scheduling fairness and QoS to keep the system responsive under load surges.
Now, async might be more efficient if you serve extreme numbers of concurrent requests from a single thread if your request processing is so simple that the scheduling cost becomes a significant portion of the processing time.
... but if your request processing happens in Python, that's not the case. Your own scheduler implementation (your event loop) will likely also end up eating some resources (remember, you're not bypassing anything, just duplicating functionality), and is very unlikely to be as smart or as fair as that of the kernel. It's probably also entirely unable to do parallel processing.
And this is all before we get into the details of how you easily end up fighting against the scheduler...
---
I'd guess the c++ event loop is more important than the jit?
Maybe a better comparison is quart (with eg uvicorn)
https://pgjones.gitlab.io/quart/
Or Sanic / uvloop?
I do a lot of Django and Nodejs and Django is great to sketch an app out, but I've noticed rewriting endpoints in Nodejs directly accessing postgres gets much better performance.
Just my 2c
I've tested this extensively on Linux. There is no more CPU used for threads vs epoll.
On the other hand, if you don't get the epoll imementation exactly right, you may end up with many spurious calls. E.g. simply reading slow data from a socket in golang on Linux incurs considerable overhead: a first read that is short, another read that returns EWOULDBLOCK, and then a syscall to re-arm the epoll. With OS threads, that is just a single call, where the next call blocks and eventually returns new data.
Edit: one thing I haven't considered when testing is garbage collection. I'm absolutely convinced that up to 10k connections, threads or async doesn't matter, in C or Rust. But it may be much harder to do GC over 10k stacks than over 8.
This seems to also be a major behind io_uring.
The issue is memory usage, which OS threads take a lot of.
Would userland scheduling be more CPU efficient? Sure, probably in many cases. But I don't think that's the problem with handling many thousands of concurrent requests today.
Co-routines are not necessarily faster than threads, but they yield to a performance improvement if one has to spin thousands of them : they have less creation overhead and consume less RAM.
This hardly matters when spinning up a few thousand threads. Only memory that's actually used is committed, one 4k page at a time. What is 10MB these days? And that is main memory, while it's much more interesting what fits in cache. At that point it doesn't matter if your data is in heap objects or on a stack.
Add to that the fact that Python stacks are mostly on the heap, the real stack growing only due to nested calls in extensions. It's rare for a stack in Python to exceed 4k.
Literally the first thing any concurrency course starts with in the very first lesson is that scheduling and context overhead are not negligible. Is it so hard to expect our professionals to know basic principles of what they are dealing with?
This is because when they are first shown it, the examples are faster, effectively at least, because the get given jobs done in less wallclock time due to reduced blocking.
They learn that but often don't get told (or work out themselves) that in many cases the difference is so small as to be unmeasurable or in other circumstances the can be negative effects (overheads others have already mentioned in the framework, more things waiting on RAM with a part processed working day which could lead to thrashing in a low memory situation, greater concurrent load on other services such as a database and the IO system it depends upon, etc).
As a slightly of-the-topic-of-async example, back when multi-core processing was first becoming cheap enough that it was not just affordable at give but the default option, I had great trouble trying to explain to a colleague why two IO intensive database processes he was running were so much slower than when I'd shown him the same process (I'd run them sequentially). He was absolutely fixated on the idea that his four cores should make concurrency the faster option, I couldn't get through that in this case the flapping heads on the drives of the time were the bottleneck and the CPU would be practically idle no matter how many cores it had while the bottleneck was elsewhere.
Some people learn the simple message (async can handle some loads much more efficiently) as an absolute (async is more efficient) and don't consider at all that the situation may be far more nuanced.
You mean concurrent tasks in the same process?
I do.
And I don't think I'm alone nor being unreasonable.
This is a quintessential example of not seeing the forest for the trees.
The point of coroutines is absolutely to make my code execute faster. If a completely I/O-bound application sits idle while it waits for I/O, I don't care and I should not care because there's no business value in using those wasted cycles. The only case where coroutines are relevant is when the application isn't completely I/O bound; the only case where coroutines are relevant is when they make your code execute faster.
It's been well-known for a long time that the majority of processes in (for example) a webserver, are I/O bound, but there are enough exceptions to that rule that we need a solution to situations where the process is bound by something else, i.e. CPU. The classic solution to this problem is to send off CPU-bound processes to a worker over a message queue, but that involves significant overhead. So if we assume that there's no downside to making everything asynchronous, then it makes sense to do that--it's not faster for the I/O bound cases, but it's not slower either, and in the minority-but-not-rare CPU-bound case, it gets us a big performance boost.
What this test is doing is challenging the assumption that there's no downside to making everything asynchronous.
In context, I tend to agree with the conclusion that there are downsides. However, those downsides certainly don't apply to every project, and when they do, there may be a way around them. The only lesson we can draw from this is that gaining benefit from coroutines isn't guaranteed or trivial, but there is much more compelling evidence for that out there.
I think rather the point is to make your APPLICATION either finish in less time, or to not take MORE time when given more load.
The code runs as fast as it runs, coroutines notwithstanding.
> I think rather the point is to make your APPLICATION either finish in less time, or to not take MORE time when given more load.
Potato potato.
And from my perspective, I don't think it's unreasonable for me to expect you to try to understand what I'm trying to communicate, rather than attempting to force me to use different words. The burden of communication is shared by both speaker and listener.
The surprising conclusion of the article is that on a realistic scenario, the async web frameworks will ouput less requests/sec than the sync ones.
I'm very familiar with Python concurrency paradigms, and I wasn't expecting that at all.
Add to that zzzeek's article (the guy wrote SQLA...) stating async is also slower for db access, this makes async less and less appealing, given the additional complexity it adds.
Now appart from doing a crawler, or needing to support websockets, I find hard to justify asyncio. In fact, with David Beasley hinting that you probably can get away with spawning a 1000 threads, it raises more doubts.
The whole point of async was that, at least when dealing with a lot of concurrent I/O, it would be a win compared to threads+multiprocessing. If just by cranking the number of sync workers you get better results for less complexity, this is bad.
Is the price of the context switching too high, or are you compensating the weakness of each system, by handling I/O concurrently in async, but smoothing the blocking code outside of the await thanks to threads?
Making a _clean_ benchmark for would it be really hard, though.
The author of "black" suggested that the cause of the slow down may be that asyncio actually starved postgres for resources:
but threads get you the same thing with much less overhead. this is what benchmarks like this one and my own continue to confirm.
People often are afraid of threads in Python because "the GIL!" But the GIL does not block on IO. I think programmers reflexively reaching for Tornado or whatever don't really understand the details of how this all works.
That is not true, at least not in general, the whole point of using continuations for async I/O is to avoid the overhead of using threads, the scheduler overhead, the cost of saving and restoring the processor state when switching tasks, the per thread stack space, and so on.
The use of async either as callbacks, or user threads, or coroutines, is a convenience layer for structuring your code. As I understand, that layer does add some overhead, because it captures an environment, and has to later restore it.
Async and parallel always use more CPU cycles than sequential. There is no question. He real questions are: do you have cycles to burn, will doing so brings the wall clock time down, and is it worth the complexity of doing so?
I was thinking this would be about using multiprocessing to fire off two or more background tasks, then handle the results together once they all completed. If the background tasks had a large enough duration, then yeah, doing them in parallel would overcome the overhead of creating the processes and the overall time would be reduced (it would be "faster"). I thought this post would be a "measure everything!" one, after they realized for their workload they didn't overcome that overhead and async wasn't faster.
Upon what the post was about, my response was more like "...duh".
Waiting for I/O does usually not waste any CPU cycles, the thread is not spinning in a loop waiting for a response, the operating system will just not schedule the thread until the I/O request completed.
You are making dinner. You start to boil water for the potatoes. While that happens, you prepare the beef. Async.
You and your girlfriend are making dinner. You do the potatoes, she does the beef. Parallel.
You can perhaps see how you could have asynchronous and parallel execution at the same time.
In the context of a Web server, a request is handled by a single Python process (so don’t give me that “OS scheduler can do other things”). Async matters here because your request turnover can be higher, even if the requests/sec remains the same.
In the cooking example, each request gets a single cook. If that cook is able to do things asynchronously, he will finish a single meal faster.
If it were only parallel, you could have more cooks - because they would be less demanding - but they would each be slower.
There is a bit of nuance here, in that the async-chef would make any individual meal slower than a sync-chef, once the number of outstanding requests is large. The sync-chef would indeed have overall higher wait times, but each meal would process just as fast as normal (eg. more like a checkout line at a grocery store).
I prefer the grocery store checkout line metaphor for this reason. If a single clerk was "async" and checking out multiple people at once, all the people in a line would have an average wait time of roughly the same for a small line size. A "sync" clerk would have a longer line with people overall waiting longer, but each individual checkout would take the same amount of time once the customer managed to reached the clerk.
This is pertinent when considering the resources utilized during the job. If an sync clerk only ever holds a single database connection, while an async clerk holds one for every customer they try to check out at the same time, the sync clerk will be far more friendly to the database (but less friendly to the customers, when there aren't too many customers at once).
The sync chef doesn't occupy the frying pan when he's boiling potatoes, so in some sense he only really does as much as he can. Having hundreds of sync chefs would likely be more efficient in terms of order volume, _but not order latency._
It depends on what you mean by "faster". HTTP requests are IO bound, thus it is to be expected that the throughout of a IO bound service benefits from a technology that prevents your process from sitting idle while waiting for IO.
Thus it's surprising that Python's async code performs worse, not better, in both throughput and latency.
> When you're dealing with external REST APIs that take multiple seconds to respond, then the async version is substantially "faster"
The findings reported in the blog post you're commenting are the exact opposite of your claim: Python's async performs worse than it's sync counterpart.
“Faster” is misleading because the speed improvements that you get with async is very dependent on load. At low levels there is going to typically be negligible or no speed gains, but at higher levels the benefit will be incredibly obvious.
The one caveat to this is cases where async allows you to run two requests in parallel, rather than sequentially. I would argue that this is less about async than it is about concurrency, and how async work can make some concurrent work loads more ergonomic to program.
> “Faster” is misleading
and
> "At low levels there is going to typically be negligible or no speed gains, but at higher levels the benefit will be incredibly obvious."
there are no "speed" gains period. the same amount of work will be accomplished in the same amount of time with threads or async. async makes it more memory efficient to have a huge number of clients waiting concurrently for results on slow services, but all of those clients walking off with their data will not be reached "faster" than with threads.
the reason that asyncio advocates say that asyncio is "faster" is based on the notion that the OS thread scheduler is slow, and that async context switches are some combination of less frequent and more efficient such that async is faster. This may be the case for other languages but for Python's async implementations it is not the case, and benchmarks continue to show this.
You are not waiting for that 1000ms, and you haven't been for 35 years since the first os's starting feature preemptive multitasking.
When you wait on a socket, the OS will remove you from the CPU and place someone who is not waiting. When data is ready, you are placed back. You aren't wasting the CPU cycles waiting, only the ones the OS needs to save your state.
Actually standing there and waiting on the socket is not a thing people have done for a long time.
The point is that async IO allows your own process/thread to progress while waiting for IO. Preemptive multitasking just assigns the CPU to something else while waiting, which is good for the box as a whole, but not necessarily productive for that one process (unless it is multithreaded).
This doesn’t surprise me at all, as I’ve had to deal with async python in production, and it was a performance and reliability nightmare compared to the async Java and C++ it interacted with.
...with the goal of making your application faster.
In doing that, you're removing natural parallelism, and end up competing with the kernel scheduler, both in performance and in scheduling decisions.
That doesn't matter though. If you think the average python user is looking for "concurrency without parallelism" with no speed/performance goal in mind, you totally have the wrong demographic.
The fact that the language chose to implement asyncio on a single thread (again the end user doesn't care that this is the case, it could have been thread/core abstraction like goroutines), with little gain, which lead to a huge fragmentation of its library ecosystem is bad. Even worse that it was done in 2018. Doesn't matter how smart you are about the internals.
Python implements things on a single thread due to language restrictions (or rather, reference implementation restrictions), as the GIL as always disallows parallel interpreter access, so multiple Python threads serve little purpose other than waiting for sync I/O. It's been many years since I followed Python development, but back then all GIL removal work had unfortunately come to a halt...
> ... no. With the goal
I assumed those meant the end user of the language (it is fair to assume the person you responded to meant that). The goal of the language itself was probably to stay trendy - e.g. JS/Golang/Nim/Rust/etc had decent async stories, where python didn't. Python needed async syntax support as the threading and multiprocessing interfaces were clunky compared to others in the space. What they ended up with arguably isn't good.
I'm pretty familiar with those restrictions which is why I expected this thread to be more of "yeah it sucks that its slower" instead of pulling the "coroutines don't technically make anything faster per se" argument which is distracting.
Then it was people saying “Guys, stop buying surgical masks, The science says they don’t work it’s like putting a rag over your mouth.”
All of these so called expert know it alls were wrong and now we have another expert on asynchronous python telling us he knows better and he’s not surprised. No dude your just another guy on the internet pretending he’s a know it all.
If you are any good, you’ll realize that nodejs will beat the flask implementation any day of the week and the nodejs model is exactly identical to the python async model. Nodejs blew everything out of the water, and it showed that asynchronous single threaded code was better for exactly the test this benchmark is running.
It’s not obvious at all. Why is the node framework faster then python async? Why can’t python async beat python sync when node can do it easily? What is the specific flaw within python itself that is causing this? Don’t answer that question because you don’t actually know man. Just do what you always do and wait for a well intentioned humble person to run a benchmark then comment on it with your elitist know it all attitude claiming your not surprised.
Is there a word for these types of people? They are all over the internet. If we invent a label maybe they’ll start becoming self aware and start acting more down to earth.
Node's JIT comes from a web browser's javascript implementation used by billions of people. It's also had async baked in from day one.
Python started single process, added threading, and then bolted async on top of that. And CPython is a pretty straight interpreter.
A comparison between Node and PyPy would be more informative, but PyPy has a far less mature JIT and still has to deal with Python's dynamism.
> If we invent a label maybe they’ll start becoming self aware and start acting more down to earth.
You can't lecture people into self-awareness, any more than experts can lecture everyone into wearing masks.
I expect upping this number would have a positive effect on asyncio numbers because the only thing[3] this[4] is[5] measuring[6] is how many database connections you have, and is about as far from a realistic workload as you can get.
Change your app to make 3 parallel requests to httpbin, collect the responses and insert them into the database. That's an actually realistic asyncio workload rather than a single DB query on a very contested pool. I'd be very interested to see how sync frameworks fare with that.
1. https://github.com/calpaterson/python-web-perf/blob/master/a...
2. https://github.com/calpaterson/python-web-perf/blob/master/s...
3. https://github.com/calpaterson/python-web-perf/blob/master/a...
4. https://github.com/calpaterson/python-web-perf/blob/master/a...
5. https://github.com/calpaterson/python-web-perf/blob/master/a...
6. https://github.com/calpaterson/python-web-perf/blob/master/a...
Sidenote here: one thing I found but didn't mention (the reason I put in the pooling, both in Python and pgbouncer) is that otherwise, under load, the async implementions would flood postgres with open connections and everything would just break down.
I think making a database query and responding with JSON is a very realistic workload. I've coded that up many times. Changing it to make requests to other things (mimicking a microservice architecture) is also interesting and if you did that I'd be interested to read your write up.
If the system as a whole is well saturated, and the python processes dominate the system load with a DB load proportional to the requests served, then I don't think we would hit any external bottlenecks.
The benchmarks performed are not that great (e.g., virtualized, same machine for all components, etc.), but I don't think the errors are enough to throw off the result. Note, of course, that such results are not universal, and some loads might perform better async.
Doesn't this prove that async is waiting for connections when you put a limit on it? The only way async wins is if it is free to hit the db whenever it needs to.
And the reasoning is explained in the article:
"The rule I used for deciding on what the optimal number of worker processes was is simple: for each framework I started at a single worker and increased the worker count successively until performance got worse."
I don't see how that is a more "realistic" asyncio workload.
It might be a workload that async is better suited for, but the point of the article is to compare async web frameworks, which will often be used just to fetch and return some data from the db.
If you had an endpoint which needed to fetch 3 items from httpbin and insert them in the db it may make sense to use asyncio tools for that, even within the context of a web app running under a sync framework+server like Falcon+Gunicorn.
In my experience Python web apps (Django!) often spend surprisingly little time waiting on the db to return results, and relatively a large amount of time cpu-bound instantiating ORM model instances from the db data, then transforming those instances back into primitive types that can be serialized to JSON in an HTTP response. In that context I am not surprised if sync server with more processes is performing better. In this test it's not even that bad... the 'ORM' seems to be returning just a tuple which is transformed to a dict and then serialized.
until then I'd had pretty much bought the hype that the new async frameworks running on Uvicorn were the way to go
I'm very glad to see this kind of comparative test being made, it's very useful, even if it later gets refined and added to and the results more nuanced
Highly disagree as the database is just another IO connection to a server, which is asyncio bread and butter. Being able to stream data from longer running queries without buffering and whilst serving other requests (and making other queries) is really quite powerful.
But yeah, if you're maxing out your database with sync code then async isn't going to make it magically go faster.
As the blog post apparently cites as well (woo!), I've written about the myth of "async == speed" some years ago here and my conclusions were identical.
https://techspot.zzzeek.org/2015/02/15/asynchronous-python-a...
It's a difficult myth to dispel and I think the situation in terms of public mindshare is much worse now than it was in 2015. Some very silly claims from the async crowd now have basically widespread credence. I think one of the root causes is that people are sometimes very woolly about how multi-processing works. One of the others is that I think it's easy to make the conceptual mistake of 1 sync workers = 1 async worker and do a comparison that way
One of my worries is that right now it feels like everything in Python is being rewritten in asyncio and the balkanisation of the community could well be more problematic than 2 vs 3.
this is exactly why the issue is so concerning for me as well.
Ok in 2015 it was a pain but with Python 3.8 it's actually a only joy & fun in my opinion.
> the balkanisation of the community could well be more problematic than 2 vs 3
If you could call Python2 code from Python3 or vice-versa as easily as you can do with async then it would be comparable.
Python packaging is something that I have fully automated (maintaining over 50 packages here) and that I'm pretty happy with.
I fail to see the problem with Python packaging, maybe because I have an aggressive continuous integration practice ? (always integrate upstream changes, contribute to dependencies that I need, and when I'm not doing TDD it's only because I have not yet proof that the code I'm writing is not actually going to be useful) That's not something everybody wants to do (I don't understand their reasoning though).
People would rather freeze their dependencies and then cry because upgrading is a lot of work, instead of upgrading at the rhythm of upstream releases. If other packages managers or other languages have packaging features that encourages what I consider to be non-continuous integration then good for them, but that's not how a hacker like me wants to work, being able to "ignore upstream releases" is not a good feature, it made me a sad developer really, "ignoring non-latest releases" have made me a really happy developer.
Most performance issues are not imputable to the language. If they are, it's probably not affecting all your features, you can still rewrite the feature that Python is not well performing for into a compiled language. I need most of my code to be easy to manipulate, and very little of it to actually outperform Python.
I've recently re-assessed if I should keep going with Python for another 10 years, tried a bunch of languages, frameworks, at the end of the month I still wanted a language that easy to manipulate with basic text tools, that's sufficiently easy so that I can onboard junior collegues on my tools, that provides sufficiently advanced OOP because I find it efficient to structure and reuse code.
Python does what it claims, it solves a basic human-computer problem, let's face it: it's here to stay and shine, and its wide ecosystem seems like a solid proof. Wether it makes sense to invest in a project or not should not depend in the language anyway.
In general, you can get higher throughput with asyncio because you don't have context switches, but it comes at the cost of latency. So hand-wavy, indeed. It really depends what sort of speed you're after.
Imagine you're loading a profile page on some social networking site. You fetch the user's basic info, and then the information for N photos, and then from each photo the top 2 comments, and for each comment the profile pic of the commentor. You can't just fetch all this in one shot because there's data dependencies. So you start fetching with blocking IO, but that makes your wait time for this request proportional to the number of fetches, which might be large.
So instead, you ideally want your wait to be proportional to the depth of your dependency tree. But composing all these fetches that way is hard without the right abstraction. You can cobble it together with callbacks but it gets hairy fast.
So (outside of extreme scenarios) it's not really about whether async is abstractly faster than sync. It's about how real developers would solve the same problem with/without async.
(Source: I worked on product infrastructure in this area for many years at FB)
At least in Typescript nowadays, the ability to just mark a function `async` and throw an `await` in front of its invocation drastically lowers the barrier to moving something from blocking to non-blocking. In the same cases if I had to recommend the same change with thread pools and callbacks (and the manual book-keeping around all that) most developers just wouldn't bother.
Yeah, that's an extremely painful way to write threaded code. Much more normal is to simply block your thread while waiting for others to .Join() and return their results, likely behind an abstraction layer like a Future.
The only time you really need to use callbacks is when you need to blend async and threaded code, and you aren't able to block your current thread (e.g. Android main thread + any thread use is an example of this). But there are much much easier ways to deal with that if you need to do it a lot - put your primary logic in a different, blockable thread.
That's not how it works. `async` and `await` are merely syntactic sugar around callbacks. Everything in javascript is already nonblocking[1], whether or not you use async/await.
[1] There are a few rare exceptions in node js (functions suffixed with "Sync"), but in the same vein, they are blocking whether or not you use async/await.
const a = an async operation
const b = another async operation
// Resolve a and b concurrently
const [x, y] = await Promise.all([a, b])
// Do something with x and y
You can naturally achieve that with callbacks but there's more boilerplate involved. I'm not familiar with Python so I don't know how it would look like without async.Edit: I just re-read your comment and the one you were responding to, and do agree that async/await don't "move" things from blocking to non-blocking. It just helps using already non-blocking resources more easily. It will not help you if you're trying to make a large numerical computation asynchronous, for example. In this regard it's very different from Golang's `go`, which will run the computation in a separate goroutine, which itself will run concurrently (with Go's scheduler deciding when to yield), and in parallel if the environment allows it.
A more interesting example would be a request that requires multiple blocking operations (database queries, syscalls, etc.). You could do something like:
# Non-concurrent approach
def handle_request(request):
a = get_row_1()
b = get_row_2()
c = get_row_3()
return render_json(a, b, c)
# asyncio approach
async def handle_request(request):
a, b, c = await asyncio.gather(
get_row_1(),
get_row_2(),
get_row_3())
return render_json(a, b, c)
# Naive threading approach
def handle_request(request):
a_q = queue.SimpleQueue()
t1 = threading.Thread(target=get_row_1(a_q))
t1.start()
b_q = queue.SimpleQueue()
t2 = threading.Thread(target=get_row_2(b_q))
t2.start()
c_q = queue.SimpleQueue()
t3 = threading.Thread(target=get_row_3(c_q))
t3.start()
t1.join()
t2.join()
t3.join()
return render_json(a_q.get(), b_q.get(), c_q.get())
# concurrent.futures with a ThreadPoolExecutor
def handle_request(request, thread_pool):
a = thread_pool.submit(get_row_1())
b = thread_pool.submit(get_row_2())
c = thread_pool.submit(get_row_3())
return render_json(a.result(), b.result(), c.result())
These examples demonstrate what people find appealing about asyncio, and would also tell you more about how choice of concurrency strategy affects response time for each request.Therefore, I would have liked to see how much memory all those workers use, and how many concurrent connections they can handle.
The underlying issue with python is that it does not support threading well (due to the global interpreter lock) and mostly handles concurrency by forking processes instead. The traditional way of improving throughput is having more processes, which is expensive (e.g. you need more memory). This is a common pattern with other languages like ruby, php, etc.
Other languages use green threads / co-routines to implement async behavior and enable a single thread to handle multiple connections. On paper this should work in python as well except it has a few bottlenecks that the article outlines that result in throughput being somewhat worse than multi process & synchronous versions.
Taken from Stephen Cleary's SO answer on this topic: https://stackoverflow.com/a/31192718
Memory is cheap; the cost is in constant de/serialization. Same with "just rewrite the hotspots in C!"-style advice; de/serialization can easily eat anything you saved by multiprocessing/rewriting. Python is a deceivingly hard language, and a lot of this is a direct result of the "all of CPython is the public C-extension interface!" design decision (significant limitations on optimizations => heavy dependency on C-extensions for anything remotely performance sensitive => package management has to deal extensively with the nightmare that is C packaging => no meaningful cross-platform artifacts or cross compilation => etc).
What? What makes you say that? What did you think I was talking about if not a production system? To be clear, we're talking about the overhead of single-digit additional python interpreters unless I'm misunderstanding something...
maybe author is concerned that many people are jumping the gun on async-await before we all fully understand why we need it at all. and that's true. but that paradigm was introduced (borrowed) to solve a completely different issue.
i would love to see how many concurrent connections those sync processes handle.
Maybe what you're getting at is cases where there are a large number of (fairly sleepy) open connections? Eg for push updates and other websockety things. I didn't test that I'm afraid. The state of the art there seems to be using async and I think that's a broadly appropriate usage though that is generally not very performance sensitive code except that you try to do as little as possible in your connection manager code.
I use trio/asyncio to more easily write correct complex concurrent code when performance doesn't matter. See "The Problem with Threads"[1].
For this use case, Async Python probably still isn't faster, but that doesn't matter. Let's not throw out the baby with the bathwater :)
[1] https://www2.eecs.berkeley.edu/Pubs/TechRpts/2006/EECS-2006-...
This is great for react or vue front end applications which get their state updated when things happen in the outside world (e.g. somebody else starts the music player, that gets related)
When CPU performance is an issue (say generate a weather video from frames) you want to offload that into another process or thread, but it is an easy programming style if correctness matters.
Concurrency can help you separate out logic that is often commingled in non-concurrent code, but doesn't need to be. As a real-world example, I used to do safety critical systems for aircraft. The linear, non-concurrent version, included a main loop that basically executed a couple dozen functions. Each function may or may not have dependencies on the other functions, so information was passed between them over multiple passes through this main loop (as their order was fixed) using shared memory.
A similar project had about a dozen processes, each running concurrently. There was no speed improvement, but the connection between each activity was handled via channels (equivalent in theory to Go's channels, less like Erlang's mailboxes as the channels could be shared). We knew it was correct because each process was a simple state machine, separated cleanly from all other state machines.
The second system's code was much simpler, there was no juggling (in our code) of the state of the system, compared to managing the non-concurrent logic. If a channel had data to be acted on, the process continued, otherwise it waited. Very simple. And it turns out that many systems can be modeled in a similar fashion (IME). Of course, we had a very straightforward communication mechanism (again, essentially the same as Go channels except it was a library written in, as I recall, Ada by whoever made the host OS).
I mean think about it. Whats the difference between sending message A and then message B versus sending messages A and B into a queue and letting some async process pop from it? Less complexity and guaranteed message delivery come for free in single-threaded code.
Am I wrong? What am I missing?
You'd need a good watchdog and error handling, but presumably some of that came for "free" in their environment.
Although if you take out the "free" OS support, watchdog, etc., I agree that there's likely a place between "shared memory spaghetti" and "multi-processing" that's simpler than both.
The other benefit of the concurrent design (versus the single-threaded version) was that it was actually much simpler. This was critical for our field because that system is still flying, now 12 years later, and will probably be flying for another 30-50 years. The single-threaded system was unnecessarily complex. Much of the complexity came from having to include code to handle all the state juggling between the separate tasks, since each had some dependency on each other (not a fully connected graph, but not entirely disconnected either). The concurrent design made it trivial to write something very close to the most naive version possible, where waiting was something that only happened when external input was needed. So the coordination between each task just fell out naturally.
You still have to care about locking the system up, but in our case because each process was sufficiently reduce to its essentials, this was easy to evaluate and reason about.
Given how slow I/O operations are, and how much modern code depends on the network, we typically need some concurrency in our code. So for me, almost always, the question isn't, "which concurrency choice is fastest?" but rather, "which concurrency choice is fast enough while leading to code with the least bugs?"
It's like multi-threading 2+2.
I suspect that the best async is that supported by the server OS, and the more efficiently a language/compiler/linker integrates with that, the better. JIT/interpreted languages introduce new dimensions that I have not experienced.
I do have some prior art in optimizing libraries, though. In particular, image processing libraries in C++. My opinion is that optimization is sort of a "black art," and async is anything but a "silver bullet." In my experience, "common sense" is often trumped by facts on the ground, and profilers are more important than careful design.
I have found that it's actually possible to have worse performance with threads, if you write in a blocking fashion, as you have the same timeline as sync, but with thread management overhead.
There are also hardware issues that come into play, like L1/2/3 caches, resource contention, look-ahead/execution pipelines and VM paging. These can have massive impact on performance, and are often only exposed by running the app in-context with a profiler. Sometimes, threading can exacerbate these issues, and wipe out any efficiency gains.
In my experience, well-behaved threaded software needs to be written, profiled and tuned, in that order. An experienced engineer can usually take care of the "low-hanging fruit," in design, but I have found that profiling tends to consistently yield surprises.
T.A.N.S.T.A.A.F.L.
While Windows has had asynchronous I/O for ages, it's still one kernel transition per operation, whereas Linux can batch these now.
I suspect that all the CPU-level security issues will eventually be resolved, but at a permanently increased overhead for all user-mode to kernel transitions. Clever new API schemes like io_uring will likely have to be the way forward.
I can imagine a future where all kernel API calls go through a ring buffer, everything is asynchronous, and most hardware devices dump their data directly into user-mode ring buffers by default without direct kernel involvement.
It's going to be an interesting new landscape of performance optimisation and language design!
> I have found that it's actually possible to have worse performance with threads, if you write in a blocking fashion
But isn't excessive blocking/synchronization not something the should already be tackled in your design instead of trying to rework it after the fact ?
I would expect profiling to mostly leads to micro-optimisations, eg combining or splitting the time a lock is taken, but when you're still designing you can look at avoiding as much need for synchronization as possible. eg: sharing data copy-on-write (not requiring locks as long as you have a reference) instead of having to lock the data when accessing it.
As another commenter says
> with asyncio we deploy a thread per worker (loop), and a worker per core. We also move cpu bound functions to a thread pool
you can't easily go from eg. thread-per-connection to a worker pool. that should have been caught during design
Yes and no. Again, I have not profiled or optimized servers or interpreted/JIT languages, so I bet there's a new ruleset.
Blocking can come from unexpected places. For example, if we use dependencies, then we don't have much control over the resources accessed by the dependency.
Sometimes, these dependencies are the OS or standard library. We would sometimes have to choose alternate system calls, as the ones we initially chose caused issues which were not exposed until the profile was run.
In my experience, the killer for us was often cache-breaking. Things like the length of the data in a variable could determine whether or not it was bounced from a register or low-level cache, and the impact could be astounding. This could lead to remedies like applying a visitor to break up a [supposedly] inconsequential temp buffer into cache-friendly bites.
Also, we sometimes had to recombine work that we had sent to threads, because that caused cache hits.
Unit testing could be useless. For example, the test images that we often used were the classic "Photo Test Diorama" variety, with a bunch of stuff crammed onto a well-lit table, with a few targets.
Then, we would run an image from a pro shooter, with a Western prairie skyline, and the lengths of some of the convolution target blocks would be different. This could sometimes cause a cache-hit, with a demotion of a buffer. This taught us to use a large pool of test images, which was sometimes quite difficult. In some cases, we actually had to use synthesized images.
Since we were working on image processing software, we were already doing this in other work, but we learned to do it in the optimization work, too.
When my team was working on C++ optimization, we had a team from Intel come in and profile our apps.
It was pretty humbling.
I think my question is whether async Python is slower in the case it was designed for -- many, long-running open sockets.
Async was traditionally used server-side for things like chat servers, where I might have millions of sockets simultaneously open.
This wasn't really the reason for the shift away from cooperative multitasking, it was really because cooperative multitasking isn't as robust or well behaved unless you have a lot of control over what tasks you have trying to run together.
In theory cooperative multitasking should have better throughput (latency is another story) because each task can yield at a point where its state is much simpler to snapshot rather than having to do things like record exact register values and handle various situations.
We've had a track record of technologies which:
1) Automated things (reliving programmers from thinking about stuff)
2) Were expected to make stuff slower
3) In reality, sped stuff up, at least in the typical case, once algorithms got smart
That's true for interpreted/dynamic languages, automated memory management/garbage collection, managed runtimes of different sorts, high-level descriptive languages like SQL, etc.
Sometimes, it took a lot of time to figure out how to do this. Interpreters started out an order-of-magnitude or more slower than compilers. It took until we had bytecode+JIT that performance roughly lined up. Then, we started doing profiling / optimization based on data about what the program was actually doing, and potentially aligning compilation to the individual users' hardware, things suddenly got a smidgeon faster than static compilers.
There is something really odd to me about the whole async thing with Python. Writing async code in Python is super-manual, and I'm constantly making decisions which ought to be abstracted away for me, and where changing the decisions later is super-expensive. I'd like to write.
None of that is true.
Even SQL modeling declarative work in the form of queries requires significant tuning all the time.
The rest of the list is egregious.
> things suddenly got a smidgeon faster than static compilers.
No, they did not.
It really didn't. Yes, in highly specialized benchmark situations, JITs sometimes manage to outperform AOT compilers, but not in the general case, where they usually lag significantly. I wrote a somewhat lengthy piece about this, Jitterdämmerung:
https://blog.metaobject.com/2015/10/jitterdammerung.html
Discussed at the time:
On the other side, you don't.
That linguistic flexibility often leads to big-O level improvements in performance which aren't well-captured in microscopic benchmarks.
If the question is whether GC will beat malloc/free when translating C code into a JIT language, then yes, it will. If the question is whether malloc/free will beat code written assuming memory will get garbage collect, it becomes more complex.
Of the things you mention, I agree on SQL, and "managed runtimes" is generic enough that I cannot really judge.
I'm thoughroghly unconvinced about the rest being faster than the alternatives (and that's why you don't see many SQL servers written in interpreted languages with garbage collection).
There's a big difference between normal code and hand-tweaked optimized code. SQL servers are extremely tuned, performant code. Short of hand-written assembly tuned to the metal, little beats hand-optimized C.
I was talking about normal apps. If I'm writing a generic database-backed web app, a machine learning system, or a video game. Most of those, when written in C, are finished once they work, or at the very most have some very basic, minimal profiling / optimization.
For most code:
1) Expressing that in a high-level system will typically give better performance than if I write it in a low-level system for V0, the stage I first get to working code (before I've profiled or optimized much). At this stage, the automated systems do better than most programmers do, at least without incredible time investments.
2) I'll be able to do algorithmic optimizations much more quickly in a high-level programming language than in C. With a reasonably time-bounded investment in time, my high-level code tends to be faster than my low-level code -- I'll have the big-O level optimizations finished in a fraction of the time, so I can do more of them.
3) My low-level code gets to be faster once I get into a very high level of hand-optimization and analysis.
Or in other words, I can design memory management better than the automated stuff, but my get-the-stuff-working level of memory management is no longer better than the automated stuff. I can design data structures and algorithms better than PostgreSQL specific to my use case, but those won't be the first ones I write (and in most cases, they'll be good enough, so I won't bother improving them). Etc.
Except for a few very special cases, it is perfectly fine to block on I/O. Operating systems have been heavily optimized to make synchronous I/O fast, and can also spare the threads to do this.
Certainly in client applications, where the amount of separate I/O that can be usefully accomplished is limited, far below any limits imposed by kernel threads.
Where it might make sense is servers with an insane number of connections, each with fairly low load, i.e. mostly idle, and even in server tasks quality of implementation appears to far outweigh whether the server is synchronous or asynchronous (see attempts to build web servers with Apple's GCD).
For lots of connections actually under load, you are going to run out of actual CPU and I/O capacity to serve those threads long before you run out of threads.
Which leaves the case of JavaScript being single threaded, which admittedly is a large special case, but no reason for other systems that are not so constrained to follow suit.
Not when you know how to call sync functions from async functions and vice versa.
An sync function can call an async function via:
loop = asyncio.new_event_loop()
result = loop.run_until_complete(asyncio.ensure_future(red(x)))
A async function can call a sync function via: loop = asyncio.get_event_loop()
result = await loop.run_in_executor(None, blue, x)
Where red and blue are defined as: async def red(x):
pass
def blue(x):
pass
Note that the documentation is wrong about recommending create_task over ensure_future. That recommendation results in more restrictive code as create_task only accepts a coroutine and not a task.This works for regular functions I don't know how it works for generators.
result = asyncio.run(red(x))
That's calling async function red from non-async code.Not seeing any particular readability issue with that usage neither. If you don't call asyncio.run or await on the result of an async function call, then you get a coroutine for result.
Still much harder to think about/read then
result = red(x)
We should be finding ways to get to the latter with concurrency. async/await is at best a patchwork compromise until we can do better.Async in this case has less ceremony than threads. But still enough to make thing explicit.
For instance, imagine a "maybe_await" method that just calls sync if is synchronous or otherwise awaits.
result = yourfunc()
if inspect.iscoroutine(result):
result = asyncio.run(result)
If I have code like that, it's at one place per library, didn't need to encapsulate it in a maybe_await function.Some Javascript frameworks, such as Vue, often do something similar in that you can pass either a sync or async callback and it does the right thing for either. In that case you could potentially inspect the function once and call it many times.
I believe this also would allow non-async function to return a coroutine I suppose.
Anyway, in this case chances are that there will be no performance overhead if there's any io bound operation running in the coroutine so it should run the iscoroutine check of the result and the await call before the function is done.
result = asyncio.run(red(x))
For example, when someone access a descriptor in Django.. this could end being a query to the db (transparent) but dangerous. With asyncio you explicitly await something to return the execution to the event loop.
At least for me sounds like a safer behaviour
But the difference between asyncio.run(red(x)) and blue(x)... isn't. There's no difference which matters. They are just different implementations of the same behaviour.
If red and blue are both the same DB query, with the only difference being red is async-style and blue sync-style, these two lines have exactly the same program behaviour:
result = asyncio.run(red(x))
and result = blue(x)
So the asyncio.run is just cognitive fog. It forces you to think about the type difference, but doesn't add any safety.It's almost the opposite of Python's usual duck-typing parsimony, which normally allows equivalent things to be used in place of each other without ceremony.
Zen is not respected by explicit asyncio, just try to compose asyncio with iterators [1]
[1] https://stackoverflow.com/questions/42448664/async-generator...
This problem doesn't exist with gevent, and composability is a desired thing in any programming language. Python's asyncio fractioned the community that was previously doing implicit asyncio with sync interfaces, and the current state of API is not an example of composable primitives that follow the Zen of Python:
> Beautiful is better than ugly.
> Simple is better than complex.
> Readability counts.
> Special cases aren't special enough to break the rules.
Meanwhile, my pretty python foo = bar().something has gone all foo = (await bar()).something
An entire category of data races like `x += 1` become impossible without you even thinking about it. And that's often worth it for something like a game server where everything is beating on the same data structures.
I don't use Python, so I guess it's less of an issue in Python since you're spawning multiple processes rather than multiple threads so you're already having to share data via something out of process like Redis and using its own synchronization guarantees.
But for example the naive Go code I tend to read in the wild always has data races here and there since people tend to never go 100% into a channel / mutex abstraction (and mutexes are hard). And that's not a snipe at Go but just a reminder of how easy it is to take things for granted when you've been writing single-threaded async code for a while.
(Not necessarily on topic, but if you’re really excited about dodging data races, I figured it would give you something fun to look at!)
[1] https://www.techempower.com/benchmarks/#section=data-r19&hw=...
As to the article the comparisons are good but fails to mention resource constraints, like Gunicorn, forking 16 instances is going to be a lot heavier on memory so for a little more RPS you're probably spending a decent chunk of change more to run your work and I don't think that's worth it considering the Async model in python is pretty easy to grok these days and under this benchmark share a similar performance profile.
Now that said If I had to guess these numbers are fine for the average API but if you're doing something like high throughput web crawling or need to serve something on the order of 10's of thousands to hundred thousands RPS async will win out on speed and resource use and ultimately cost.
Plus at one point they were like "we could only get an 18% speed up with Vibora" haven't used them my self. But 18% performance increase at really any level of load is fantastic. Hand waving that off tells me the work loads for what is "realistic" don't take in to account real high RPS workloads like you might see at major tech companies.
It really depends on how the application is designed. Fork operates through mmap and copy-on-write. It's extremely lightweight by default.
A well-designed fork-based application will already have everything necessary to run a given process into memory, not munge any of the existing shared memory, and only allocate and free memory associated with new events/connections/etc.
When programmed that way, individual forks are incredibly light on resources. All the workers are sharing the exact same core application code and logic in memory.
Oh interesting, are you saying an intelligent forking implementation is able to share static portions of memory with multiple children?
I was perhaps under the naive assumption forking was pretty much just a full memory copy of the parent.
[1] https://www.informit.com/articles/article.aspx?p=368650
Concurrency is many things at once. That's it.
Async frameworks end up with better concurrency properties because you're not paying the memory and context switching overhead of an entire 'thread' for each thing that you are trying to do at the same time. Instead you are paying the (normally cheaper) overhead of what is essentially a co-routine call.
The disadvantage being that you have to manage these context switches yourself, and that they tend to happen more frequently (to maintain the illusion that we are doing many things all at the same time on a single cpu). There is no way that an async framework would ever have better straight line performance than a synchronous one, simply because of all of these extra context switches, and that's fine because that's not what it is for.
Imagine I want to have 10,000 requests held open at the same time. Your flask server with 16 workers is going to have a tough time as you don't have enough workers to service that many threads, requests won't get serviced and things will start to time out. Because an async framework multiplexes those workers so that can each individually handle multiple requests at once. Multiplexing in this way costs you something performance wise.
If you were to crank up the concurrency beyond 100 at once (the default in the posted scripts), you would start getting different results.
If you want to actually go faster the asyncio interfaces used by aiomultiprocess module get you there by maintaining the event loop across multiple processes. You can save time and memory by sharding your data set and aggregating the return data.
So there's still utility, so YMMV.
Now the next thing I'd be interested to get debunked is multithreading vs multiple processes with shared memory (SysV shmem). I'm not very sure, but I'd not been surprised to hear that the predominance of multithreaded runtimes (JVM, most C++ appservers) is purely a cargo-cult effect. As far as I remember, threads were introduced for small and isolated problems in GUI programs, like code completion in IDEs; they were never intended for replacing O/S processes and their isolation guarantees.
Threading is faster, but really only if youre willing to give up your locks and design for it properly.
But now there are solutions both existing and upcoming such as Go and Java Project Loom that fix that one flaw. I don't see much appeal in the reactive style at this point.
Surely this is back to front. It is go faster stripes because it doesn't make it go faster.
For what it's worth, I think people are using aiopg because it works with SQLAlchemy whereas asyncpg does not.
I kept the database driver the same because I'm testing sync vs async and not database drivers. I would be interested in testing asyncpg, particularly a performance claim is a big part of that library's documentation but another time.
uvicorn --port 8001 --workers $PWPWORKERS app_starlette:app --loop uvloop
The uvicorn docs should point out what a big difference uvloop makes.
If I'm not wrong, aiopg it's something not fully async, just because relais on the old driver.
I've thoroughly benchmarked my own framework[1] for REST APIs and now it outperforms many of Go / Node.js platforms on Techempower[2]
[1] https://github.com/gotzmann/comet
[2] https://www.techempower.com/benchmarks/#section=test&runid=e...
Anyway, I'm running quite a few small Python services on cheap VPSs, i.e., shit performance, and using async was beneficial for me, with performance being ~30% better. They are bread-and-butter apps that read from Postgres, do some HTTP requests, process the results, and potentially write stuff back to the DB. Same performance gain for other services that have HTTP servers.
In my mind, hardware can be used more efficiently with async, since while one routine is waiting for an async result, other routines can run meanwhile.
> Why the worker count varies
> The rule I used for deciding on what the optimal number of worker processes was is simple: for each framework I started at a single worker and increased the worker count successively until performance got worse.
> The optimal number of workers varies between async and sync frameworks and the reasons are straightforward. Async frameworks, due to their IO concurrency, are able to saturate a single CPU with a single worker process.
> The same is not true of sync workers: when they do IO they will block until the IO is finished. Consequently they need to have enough workers to ensure that all CPU cores are always in full use when under load.
With a sync server, a worker is inactive as long as it is waiting on IO, so you need enough workers to maximize the chance that all workers are busy, else, some clients are waiting even if you've got the CPU to deal with them.
With an async server, a single worker handles many clients simultaneously, in theory, a single worker per core is sufficient to eat all the CPU available.
Hi - that is explained in some detail in the article
Does that get you anything, or am I misunderstanding how multiprocessing/fork works?
[1]: https://github.com/tiangolo/fastapi
[2]: https://github.com/tiangolo/full-stack-fastapi-postgresql
[3]: https://github.com/tiangolo/uvicorn-gunicorn-fastapi-docker
Perhaps someone here can explain what the asyncio paradigm does for you beyond "ease of use" when it doesn't get you past the single processor / GIL issue. In what environments are the "os threads" created by the Python engine actually that expensive? I suppose if you are just starting out it may be easier to grok, but then it won't transfer as well to other programming environments besides perhaps NodeJS.
For async python, when you make 1000 requests, does it immediately register 1000 jobs across your CPUs via workers for processing? Does that just mean each job takes a tiny piece (1/1000) of the resource pie resulting in slower performance for all jobs?
Whereas in sync python you are saying you can only perform X number of jobs at a time where X is the number of allocated workers. So resource allocation is roughly divided into X parts.
You also have a db connection pool layer after the server code. Isn't that ultimately your bottleneck? I wonder if your async server is saturating the CPUs making the connection pool slow.
@calpaterson can you provide guidance?
I'd like to try an alternative query pattern. The current pattern implemented in the benchmarks is select 1 row in 1 query. I'd like to try an implementation with 2 queries - select count() from table and select * from table limit 10, which is a very common pattern for a REST list view. I would hypothesize that the async apps would perform better in this case, but I'm curious what this benchmark with show.
Before you start you should know that Tudor M (see a PR on the project) experimented with changing the query patterns (to three queries, but not a count(*)). It doesn't change matters and the basic reason for that is that nothing has changed - simply having more blocking or non-blocking IO is irrelevant to throughput - except that the more yields you have the more problematic your response times are going to be under load.
They also misleadingly de-emphasise latency variation. One Python async framework I'd never heard of was top of the pops on throughput there even though the latency numbers suggested it had pretty much fallen apart in the test.
The takeaway of this article is that python’s async io implementations perform poorly.
That’s surprising, since async I/O is usually used for performance reasons, and in most other languages, async I/O can be much faster.
https://www.techempower.com/benchmarks/#section=data-r19&hw=...
16 workers is not that much considering that modern servers can have a lot of cores available, and I expect that the more workers you need the more likely you'll hit other bottlenecks:
* the more workers you need, the more memory you consume (workers are processes, not threads),
* I don't know how OS scheduler behave these days, but the general-purpose OS scheduler may consume some CPU time you'd rather give to your app.
I understand the point on latency variation though: preemptive multitasking will slice the CPU time "fairly" between workers, while in a cooperative multitasking situation, it would be the job of the programmer to yield after some time.
I suspect that scheduler overhead is not a realistic consideration for a Python program. My understanding is that switching executing process takes microseconds at worst, which would be too small to notice from the point of view of a Python programmer.
On "it would be the job of the programmer to yield after some time" - I'm always personally suspicious of any technique that rests on programmer diligence. My experience suggests not to require (or even expect!) programmer diligence, even from my own (I assure you, god like) programming abilities. Secondly, yielding more often probably would not help (and in fact I half-suspect part of the problem is the frequent yielding at every async/await keyword!).
Edit: I've been downvoted so I'll add a precision: Usually, it is believed that async shines against other models once you reach a certain scale (https://en.wikipedia.org/wiki/C10k_problem). This benchmark shows than async app frameworks are slower than the sync ones when running at a given scale, and since the article doesn't give many details on the incomming traffic, I can only assume that it's low, since it saturates 4 cores.
I believe that your conclusion that "Async python is not faster" is an over generalization of your use case.
I'm not saying that the configuration in your benchmark is not correct, I am saying that this benchmark may not yield the same results if you try to scale it on bigger hardware.
I believe that scheduler overhead can't be ruled out (not for python nor any other program) on a server since we've sometimes observed that the scheduler could be the bottleneck under some circumstances. For instance, some Linux schedulers used to show poor perfs when using nested cgroups with resources quota enabled.
Also, I'd like to state my first point again: you need to see how the number of workers will influence the memory usage on your system. Especially with python, if you've got a lot of workers, you can expect some memory fragmentation that can impact the perf of your system.
Old stories that come again and again
[1] https://discuss.ocaml.org/t/multicore-ocaml-may-2020-update/...
Now yield all those calls asynchronously as an array. What is this even about?
If one connection needs to do a lot of work (your 10 queries), then that's a different (also important) problem to solve. Async is a nice (from a programmer point of view) way of doing it.
https://github.com/nDmitry/web-benchmarks
Long story short - asyncio is twice as fast... (results are at the bottom of the readme).
I’m not familiar with python but it seems like there is a glaring performance bug iff using one thread per connection is faster than using async io.
So for the c10k problem on a machine with 2 GB of RAM, async will win because threads will exhaust the memory of the machine. Give that same machine 200 GB of RAM and threads may end up being faster.
There is also the not much discussed issue of having shared resources between all of these threads and the impact of such a threading model on the engineering part if writing such a program. I personally haven’t seen the thread per connection model in a successful large scale server.
It also would be interesting to see the memory footprints of the different solutions.
I think the Rachel by the Bay blog post does a good job of explaining what event loops are doing under the hood, and how that can lead to bad tail latency for web requests.
Increasing throughput doesn't mean faster, it means more efficient use of your resources.
Asynchronous and parallelizing workloads only increases throughput not speed. You get more done faster, you don't get each thing produced faster..er.
---
Async is only faster if you're not CPU constrained. I don't think anyone is surprised.
The following is really simplified; but, hopefully this makes things more clear for folks...
Assuming ONE cpu, with ONE thread
Synchronous call:
[A: Start]---------------->[B: Finish]
Asynchronous call:
[A: Start]-------->[B: Pause]...(sleep)...[C: Resume]----->[D: Finish]
There is no way to make the async call faster than the synchronous call, period. By simply having the operation pause/wait/resume (context switch) it has introduced overhead that is not present in the synchronous operation.
So WTF async?
Async is only useful when the context switching overhead is less than the time the I/O operation takes. That's it... So when you have I/O bound tasks, that take more time than it does to switch contexts (and carefully manage how many context you have), you can have increased _throughput_.
MxN (kernel threads to green threads) seems to be the established wisdom for go now, with other langs catching up, but unlike python, go has no GIL so is 'less likely to get stuck' (I say, without a footnote)
Wider focus on observability makes me optimistic that we'll do better as an industry at packing software onto hardware, and understanding which workloads benefit from what.
But Python is an excellent language for quick prototyping and for controlling other things, like coordinating GPUs who do the actual compute work. So I don't quite get why we need to make Python usable for Webservers, when we already have other languages optimized for that purpose, e.g. Google's Go.
uwsggi+flask - 16 workers unicorn+starlette - 5 workers
The highest throughout examples in your benchmark all have 16 workers. I also don’t see any hardware data... Does your machine have 16 cores?
One other thing I noticed is that this uses aiopg for the async db queries instead of asyncpg, which is more widely adopted and IMO much better.
I was hoping to re-run these benchmarks myself with asyncpg.
Looks like actually running the benchmarks would take a bunch of manual work. In fact I don’t see instructions for running these benchmarks to replicate your results.
How is this an apples to apples benchmark ?
No response time simulation...
Learn what async code does for you.
> async I/O is faster because it avoids context switches and amortizes kernel crossings
I think this is widely believed, but it's not particularly true for async I/O (of the coroutine kind meant by async/await in Python, NodeJS and other languages, rather than POSIX AIO).
With non-blocking-based async I/O, there are often more system calls for the same amount of I/O, compared with threaded I/O, and rarely fewer calls. It depends on the pattern of I/O how much more.
Consider: with async I/O, non-blocking read() on a socket will return -EAGAIN sometimes, then you need a second read() to get the data later, and a bit more overhead for epoll or similar. Even for files and recent syscalls like preadv2(...RWF_NOWAIT), there are at least two system calls if the file is not already in cache.
Whereas, threaded I/O usually does one system call for the same results. So one blocking read() on a socket to get the same data as the example above, one blocking preadv() to get the same file data.
Every system call is two user<->kernel transitions (entry, exit). The number of these transitions is one of the things we're talking about reducing with async/await style userspace scheduling.
Threaded I/O puts all context switches in kernel space, but these add zero user<->kernel transitions, because all the context switches happen inside an existing I/O system call.
Another way of looking at it, is async replaces every kernelspace context switche with a kernel entry/exit transition pair instead, plus a userspace context switch.
So the question becomes: Does the speed of userspace context switches plus kernel entry/exit costs for extra I/O system calls compare favourably against kernel context switches which add no extra kernel entry/exit costs.
If the kernel scheduler is fast inside the kernel, and kernel entry/exit is slow, this favours threaded I/O. If the kernel scheduler is slow even inside the kernel (which it certainly used to be in Linux!), and kernel entry/exit for I/O system calls is fast, it favours async.
This is despite userspace scheduling and context switching usually being extremely fast if done sensibly.
Everything above applies to async I/O versus threaded I/O and counting user<->kernel transitions, assuming them to be a significant cost factor.
The argument doesn't apply to async that is not being used for I/O. Non-I/O async/await is fairly common in some applictions, so that tilts the balance to userspace scheduling, but nothing precludes using a mix of scheduling methods. In fact doing blocking I/O in threads, "off to the side" of an async userspace scheduler is a common pattern.
It also doesn't apply when I/O is done without system calls. For example memory-mapped I/O to a device. Or if the program has threads communicating directly without entering the kernel. io_uring is based on this principle, and so are other mechanisms used for communicating among parallel tasks purely in userspace using shared memory, lock-free structures (urcu etc) and ringbuffers.
What’s going on here is python specific.
2. For this specific test, there is a downside to async, as shown by the test. Even if what's going on here were Python-specific (which is still something you made up), downsides to async which only occur in a Python environment are still downsides to async. The title of this post is "Async Python is not faster"--that conclusion is incorrect for many reasons, but none of those reasons include the words "NodeJS", "JS", or anything else that is not in the Python ecosystem.
3. What is going on is probably specific to the tools being used, which is why I said "those downsides certainly don't apply to every project". In fact, they probably don't apply to the idiomatic ways of implementing this in Tornado, for example. But note how I said "probably" because I don't know for sure, and I'm not comfortable with making things up and stating them as facts.
That's rude.
Let's put it this way. NodeJS and nginx leveled the playing field. It destroyed the lamp stack and made async the standard way of handling high loads of IO. From that alone it should indicate to you that there is something very wrong with how you're thinking about things.
You know the theory of asyncio? Let me restate it for you: If coroutines are basically the SAME thing as routines but with the extra ability to allow tasks to be done in parallel with IO then what does that mean?
It means that 5 async workers in theory should be more performant than 5 sync workers FOR highly concurrent IO tasks.
The logic is inescapable.
So what does it mean, if you run tests and see that 5 async workers are NOT more performant than 5 sync workers ON PYTHON exclusively? The theory of asyncio makes perfect logical sense right? So what is logically the problem here?
The problem IS PYTHON. That's a theorem derived logically. No need for evidence or data driven techniques.
There's this idea that data drives the world and you need evidence to back everything up. How many data points do you need to prove 1 + 1 = 2? Put that in your calculator 200 times and you got 200 data points. Boom data driven buzzword. That's what you're asking from me btw. A benchmark, a datapoint to prove what is already logical. Then you hilariously decided to dismiss it before i even presented it.
Look, I say what I say not from evidence, but from logic. I can derive certain issues about the system from logic. You just follow the logic I gave you above and tell me where it went wrong and why do I need some dumb data point to prove 1+1=2 to you?
There is NOTHING made up above. It is pure logic derived from the assumption of what AsyncIO is doing.
>But note how I said "probably" because I don't know for sure, and I'm not comfortable with making things up and stating them as facts.
But you seem perfectly comfortable in being rude and accusing me of making stuff up. I'm not comfortable in going around the internet and trashing other peoples theories with accusations that they are making shit up. If you disagree say it, I respect that. I don't respect the part where you're saying I'm making stuff up.
> That's rude.
It wasn't rude, it was predictive, and I predicted correctly. You literally ignored the second half of the sentence where I already explained why your incorrect conclusion is incorrect.
Your logic makes perfect sense, in a world where I/O bound processes, JIT versus interpretation differences, garbage collection versus reference counting differences, etc., don't exist. But those things do exist in the real world, so if your logic doesn't include them, you're quite likely to be wrong. In general, an interpreted concurrent system is far too complex to make performance predictions about based only on logic, because your logic can't possibly include all the relevant variables.
> No need for evidence or data driven techniques.
Well, there's where you're wrong. It turns out that if you actually collect evidence through experimentation, you'll discover results that are not predicted by your logic.
> Then you hilariously decided to dismiss it before i even presented it.
Well, you presented basically what a predicted, so... I wasn't wrong.
> But you seem perfectly comfortable in being rude and accusing me of making stuff up. I'm not comfortable in going around the internet and trashing other peoples theories with accusations that they are making shit up. If you disagree say it, I respect that. I don't respect the part where you're saying I'm making stuff up.
You are, in fact, making stuff up. If you are offended by accurate description of your behavior, behave better.
Let me give an example: If you run a web-service, chances are that you're gonna make some network call as part of processing a request, be it database, network, search index, etc. With async, these calls free up your application to work on another requests until the call returns. I'd say that many modern web-services are just a stitching-together of external calls (get-user-info-from-db, retrieve-user-items-from-db), so the only real work left for the web application to handle is to wait and parse/encode some JSON to client. If you can't max out your bandwidth in terms of JSON performance with one core you just start as many async processes as you have cores. The big advantage over sync processes then is that your async process can also handle insanely-long-running background calls while still crunching through the other requests. Someone please explain to me if I'm missing something here.
If you're writing synchronous code and not using threads, then yes, your analysis is right, but that would be a daft thing to do!
The difference between sync and async is mostly whether the state associated with a task (like an incoming HTTP request being serviced) is kept on a dedicated native thread stack (as it is with sync), or in some sort of coroutine structure (as with async). Thread stacks may be somewhat more efficient, but you have to allocate a whole stack upfront for each thread, so if you want to have lots of threads, you need to dedicate a lot of memory to that, and that goes badly. For applications with small numbers of tasks in flight at once, we shouldn't expect a lot of difference between sync and async code. But for tasks with huge numbers of tasks (chat servers are the classic example, but high-traffic webservers with lots of blocking calls in the backend are another), async code should keep chugging on where sync code just falls over.
tl;dr async is about the number of tasks you can handle at once, not the speed with which you handle each task.
You can run more queries concurrently but then the database just switches between those, which usually means worse latency and not potentially worse throughput.
As you wrote yourself an application handles request probably by doing a few database queries. As a consequence regardless of how many requests you can handle concurrently, they are probably all bottlenecked by the number of CPUs on your database server.
All of this is probably true regardless of whether you are using async or sync io. However with async io you have some overhead.
Of course, if you also happen to be doing other kinds of IO, say HTTP requests to other services in addition to or instead of those database queries. This might look a bit differently.
tl;dr: It really depends on what your application is doing.