Async Python is not faster (2020)
calpaterson.com
calpaterson.com
And a rebuttal from Python core dev @ambv at the time pointing out how the methodology's flawed: https://twitter.com/llanga/status/1271719778324025349
It was slower with the fewer threads, and giving it more threads made it even slower. Celebrating this as awesome efficiency is..."interesting".
And saying that "being fast is not the goal" doesn't debunk the result or make its methodology flawed. Quite the contrary, it raises a good discussion about what those goals may be and clears up misconceptions, because apparently many people do believe it is faster (and for many async-style APIs those claims are either made or at least strongly hinted at and not disavowed. Looking at you, GCD/libdispatch).
With 1 to 4 process increase in 4 cores there's basically no granularity, at 3 there's not enough concurrency, and at 4 there's too much, choking the PostgreSQL process.
Whereas 1 to 16 blocking process increase allow for more granularity ramp up and leave the (almost) exact room for the PSQL process to do its thing before performance degrades.
That would be solved having the DB on a separate host to isolate this effect, and what happens in literally any non-dev environment.
And that's leaving asid the fact that 16 processes use roughly 4 times as much memory as 4, meaning async is cheaper to scale horizontally overall.
Async solves the problem of not wanting to block other requests while you go out to expensive external resources. If you benchmark stuff while it's all on the same machine, it masks that cost.
I couldn't find the size of the DB. Is it 100GB of data? or more like 10MB? Doing one point query on 10MB might not touch even disk a whole lot.
If you were using srcreigh's hosted DB service and queries took 1s to run, the sync servers throughput would suffer greatly, meanwhile async servers would still perform pretty good.
The minute your db queries become expensive the sync workers end up slowing down and stopping other functions.
If your workers are tied up doing expensive queries you might be holding up cheap queries that might be getting data from a cache or whatever.
There's also this myth that Python code is free if the i/o wait is long enough. That's not true. 1). 100 queued connections on an async python server sucks so much momentum out your application. Especially in a resource constrained cloud environment. 2). Any server awaiting a database is going to transform the data to some response object which is going to wreck your overloaded async server. 3). As far as I know, all the async WSGIs are written in Python and will use WAY more CPU than the C-based sync WSGIs.
The bottle neck could be any form of downstream service, possibly one you do not control.
Not everything is a CRUD app running in the cloud.
Regardless, HTTP servers are completely uninteresting on their own. No one makes a request to a web server to see how the server is feeling. The server's only function is to serve content. Content it must retrieve (i/o), store (memory), and transform (cpu). That store and transform step nukes your RPS and is completely unavoidable. And, in a high worker environment, will crash your system.
1. https://www.youtube.com/watch?v=A-RfHC91Ewc
On a more serious note, describe the service for me and I'll give an opinion. Clock as a service is not specific enough.
Also: how did we come to Python as a web server when there are languages and applications better suited? Do we really want one language to be everything to everyone?
Python does a lot well for a lot of people. All that said, I think a lot of languages can do a lot of things well without serious performance impact. If you need the performance, choose the language based on that criteria.
That's why I use PERL for my web server. ;)
I'm not denying Python is an amazing language. I have to use it for machine learning, and I've grown to really like it since the first time I used it in the 90's.
Just seems odd. Like, if you had a catering business and you drove a Ford Escort, and decided to add a bunch of trailers to it because it had too little capacity in the back seat. And then tried to optimizes those trailers because it was too inefficient, rather than just buying a delivery truck suited for delivering food. OK, that was an awkward analogy.
But to the other poster's point, some people use what is natural: PERL is so absolutely burned into my brain after 10 years writing CAD flows with it, that when I need to munge a file I've already written the code in my head before I've even opened my text editor.
Yep, and your 16-worker gunicorn is going to serve 16 rps. What's shown here (nginx, pg, Python application running on the same low CPU VPS, one extremely fast db query per request) is not a "realistic benchmark".
However, if 10% of your requests take 10s, your 16 workers are very soon going to become 16 stuck workers and you won't be able to fulfill any new requests. This is the problem that async solves.
You shouldn't use a fire hose to water your plants but you also shouldn't use a sprinkler system to put out a fire.
One of my systems does around 8 api calls to service a request. It serves 80,000 daily users generating close to 2,000,000 expensive api calls and typically thats served by a single node running on a 2 core cpu. Gunicorn/starlette/fastapi.
The real problem with async in python is how easy it is to break it by introducing code or dependency that hogs the cpu every now and then. This usually means debugging weird timeouts that only happen every few days and are super hard to trace. Not sure I'd like to do async python again but it sure is efficient for I/O heavy workloads.
As long as I keep my async code on the webserver end, and my sync code elsewhere, the project seems to stay organized well. Otherwise, I tend to lose sight of proper naming, etc.
Async/await at the language level is a total scam. Give me Nginx+Lua and let me just write my regular procedural code and the underlying runtime will handle yielding/resuming for me.
It just so happened that we were trying to break 100k connections on a single machine at the same time, which requires you to avoid thread switching. Something which an async executor is also doing.
Those two goals got combined into to "the great next thing that's better in every way" when in reality, 99% will never hit that bottleneck, and the people who write good async code could write it in either of the two other forms as well.
That's because they both use generator cleverness under the hood. What asyncio does do is remove all the monkey patching that makes gevent work.
I used asyncio for the first time seriously recently and was pleasantly surprised. But, crucially, I didn’t want throughput, I wanted a low-cpu use app with easy-to-write concurrency.
I really hope GIL goes away and we can just use threads.
Some people are looking at ways to solve this. I know urllib3, elasticsearch-py, and a few others use unasync (https://github.com/python-trio/unasync) to transform async code into sync code, leaving one codebase supporting both uses in different namespaces. This leaves you with some conditional logic (is_async_mode() -- https://github.com/python-trio/hip/blob/master/src/ahip/util...). I'm seriously considering this approach.
oh my goodness..
So now we have sync code, that we've made into async code, that we now have to turn back into sync code! It's just workarounds upon workarounds upon workarounds. This is beyond the pale.
Dealing with threads directly is significantly worse. It's harder to do task cancellation, harder to do timeouts, you often need to reach for locks for the simplest things.
I found working with GCD was actually pretty good: of course it doesn't solve the safety concerns which rust aims to, but it's a sensible set of primitives which can be composed to achieve virtually whatever async problem you want.
I think there are two possibilities:
1. We haven't really landed on the right abstraction for async yet. Current approaches try to hide the wrong bits of complexity and this makes async difficult to reason about, or at least makes it a bit of a minefield of special cases and considerations.
2. Or, async programming is just complex, and there's never going to be a way to tie it up as a neat little consistent language feature.
I think there are two 'correct' abstractions for async. One is where you build it as a high level abstraction on top of other concurrency primitives. Elixir's Task.async* functions do this spectacularly, and even do things that you don't get natively from Erlang that make it very nice to work with (by giving you an opinionated way of keying into global state partitions).
The other "correct" async abstraction, imo, is a very thin, low-level abstraction over function frames and context switches. The zig programming language gets this right, IMO. I was able to write a feature that wraps zig's low-level async primitive in an event loop around the erlang's cooperative FFI scheduler, the rust equivalent hasn't figured that out yet, and since zig is colorless, you can use the same function in a non-async context, the yield points are just no-ops (which rust will never let you do, for reasons beyond just safety).
Everything a modern computer does has potential latency - whether that's waiting for the database to respond or waiting for L2 cache to respond. The programmer and the application needs decide when it's worth waiting around for the result, and when it's important to do other work in the meantime.
What about Go?
As for generics, they're going to be added in the next release of Go, scheduled for February 2022.
This view will magically become mainstream as soon as "having worked with async" stops being a resume inflating selling point. I'm very glad Java is doing the right thing and hiding everything under the same "Thread" abstraction.
AFAIU async enables greater scalability by allowing computation to continue in other contexts while awaiting read/write operations - at the cost of slightly lower single context performance.
Judging from what I've read and seen, yes. Async may scale better, but are you working at scales where that pays off?
The article, and rjbwork, use slower to mean "slower when the CPU isn't regularly idling". In this context, Node.js async is slower as well, that's why the library better-sqlite3 for NodeJS is synchronous (because you're rarely going to be using sqlite in a context where the CPU/memory isn't very busy).
I may have read the article a bit differently, it sounded like "don't bother too much with async, it's slower, even for realistic web servers".
In your I/O-bound Node.js server, if you would exchange all async calls with synchronous ones, your performance would get a big hit.
I don't think anyone would argue that async code makes CPU-bound computations faster, the opposite is the case because you have naturally more overhead.
So, it's not a problem when the thread sleeps while waiting for the end of the I/O operation, but it's not optimal.
The linux kernel has a pretty good handle on multitasking, blocking, and scheduling so I don’t see the desire to rebuild the wheel in userspace with coroutines to achieve what the kernel already does.
[1]: https://serverfault.com/questions/1012360/how-can-i-tell-the....
I guess I should be using a normal synchronous framework that just calls asyncio.run within each api call?
So if you have some server and you want to handle hundreds of thousands of concurrent requests, and each requests takes a few hundred milliseconds and most of this time is waiting for I/O (e.g, waiting for other micro services to reply), then using Async IO you could do this in one single box. Yo probably won't be able to do that with a typical balanced process pool such ash uwsgi or gunicorn.
Yes and no matter how many times its said people will still opt for the async option even though its slower in the general case. I went so far as to benchmark this myself in 2019 using EC2, digital ocean, and my local machine. I nearly published a blog post but I don't appreciate the attention that brings.
Gunicorn is not stable under heavy load. uWSGI delivers better throughput. If we talk about cloud environments exclusively, uWSGI is better full stop.
If you have endpoints with long awaits consider a strategy other than holding open the connection. If you need to run a socket server then run it separately from your api server.
Daphne is written to be a reference implementation of ASGI, not as a Django-specific server.
If you're using Django Channels in production you might even be better served by trying out Uvicorn, it's a lot faster.
Actually, I found that in some places the code wasn't perfectly "async", or there were some hidden CPU intensive parts, so that every now and then the GUI still had responsiveness issues. I then moved all the network code to a separate thread. Then I had two event loops, one from the toolkit (Gtk) on the main thread, and one from Twisted on the background thread. For communication, I had functions to "submit" a function to run on either thread (a form of message passing). Personally, I found this to be the best way to architecture a complex GUI app. For such an app, the most imporant things are that it is responsive and correct, so this works pretty well.
- async python is faster, but for a very specific niche of workload. Most tasks you do don't fit that workload at all. Which mean most of the time, you should NOT use async Python. And it's ok. Don't make your code complicated when you don't need it. Chose a tech for the need, not the hype.
- async python is not just for performances. It helps with making some specific kind of concurrency easy to reason about, because the context switch is explicit, and the chain of event can appear linearly in the code, thanks to await.
- it's ok to have some sync processes and some async processes. It's not one or the other.
- async is not a replacement for threads or processes. In fact, there are good use cases for having several processes, each with several threads, each with one event loop. Which also means that if you benchmark your WSGI code with 16 workers, you should do so as well with your aWSGI one.
The corollary to this is that some web site loads are very well suited for async (e.g: an SPA with a lot of connectivity which delegates long running code to other services), but a lot are not. If each request makes a long SQL query, seeks the hard drive, dynamically performs i18n and adds some calculation on top, it may very well block the event loop for too long, killing any benefit.
So when would you use async python ?
- For performances, when you need to maintain numerous long connections. E.G: doing websockets ? Use asyncio. Serving static files without nginx ? Use asyncio. Want to create a web crawler and your memory budget does not allow to open 10000 threads ? Use asyncio.
- For the interface, use the async/await keywords everywhere you need need inversion of control for I/O. It's not just about perf here, it's also a mechanism to delegate arbitrary parts of your code with a common standard interface around a fancy state machine + scheduler. You can use that to abstract all your I/O and switch backends at will, while offering callback inlining.
But again, you have awesome threading and multiprocessing pools with python. Not to mention tools like zeromq, scrapy or celery, which do a lot for you. Don't run to async just because "it's faster".
Having async by default in frameworks like fastpi opens up a ton of possibility though. Live settings, pub/sub between processes, websockets...
Assuming you're disagreeing with this, yeah, I've seeing "easier concurrency" be a massive footgun because you have code that's safe as long as an await doesn't get added in the wrong place.
I would say it's pretty hard to add an await in the wrong place, although not impossible. I think it's easier to:
- forget an await
- forget a asyncio.wait or a asyncio.gather
The later is very common, and is being addressed by the structured concurrency movement we've seen with Trio and the likes. This should end up in mainline python eventually because it is, indeed, a problem.
However, it's still way easier to shot yourself in the foot with concurrency and threads than with asyncio.
However, it is harder to startup with asyncio than with threads, because there is more to learn just to boot.
Do you use a typechecker?
It worked really, really, really well.
I'm not saying I'd rather use Async over web workers, but when you have the right task for this tool and you architect the solution around it intelligently it really does shine.
Async should generally be better for throughput because you can complete work items without interruption, whereas context switching will give all threads a chance to make progress. Async will start to have issues once you do any sort of non-trivial processing in the event loop. The irony is most cases employing async are latency-sensitive, not throughput-sensitive.
> Uses a VPS
Not saying the results are wrong (though for client side code my experience is very different), but that's not really a good start.