How Python asyncio works: recreating it from scratch
jacobpadilla.com
jacobpadilla.com
The biggest problem with asyncio is how easy/common it is in Python to be able to block the asyncio thread with synchronous calls, gumming up the whole system. Python sorely needs a static analysis tool that can build a call graph to help detect if a known thread-blocking call is called directly or directly from an async def.
I have toyed with a MyPy plugin to do this, but the arbitrary state storage/cache is a bit limited to do good call graph and function coloring (metadata can be on types but not functions). And there's just not many other good _extensible_ static analysis libraries. Language like Go have static analysis libraries that let you store/cache arbitrary state/facts on all constructs.
When I found out how it had implemented the asyncio event loop it was a real mind blown moment.
Forgive my ignorance, but I always figured Python just doesn't scale well for this use case.
Although top positions do a lot of "cheating", take a look where most Python entries sit: https://www.techempower.com/benchmarks/#hw=ph&test=fortune&s...
Fastest Python entry starts at 234th place while .NET one is at 16th (and that with still somewhat heavy impl., upper bracket of Fortunes is bottle-necked by DB driver as well which is what top entries win with).
I had the same feeling here, .net is just faster , but I think a lot of teams get into this workflow where they already have a Python solution and they don't want to rewrite it.
I'm not saying everything needs to be written in a systems language like Rust, but Python always strikes me as a weird choice where performance is a concern.
I'm pretty amazed to see a JavaScript runtime ranking so high here
I think an argument could be made that without real life use cases, these metrics are useless.
Comparing purely interpreted languages or interpreted + weak JIT compiler + dynamically typed (Python, JavaScript, Ruby, PHP, BEAM family) to the ones that get compiled in their entirety to machine code (C#, Java, Go, C/C++/Rust, etc.) is not rocket science.
There is a hard ceiling as to how fast an interpreter can go - it has to parse text (if it's purely scripting language) first and then interpret AST, or it has to interpret bytecode, but, either way, it means spending dozens, hundreds or even thousands of CPU cycles in a worst case scenario per each individual operation.
Consider addition, it can be encoded in bytecode as, let's say, single 0x20 byte followed by two numeric literals, each taking 4 bytes (i32). In order to execute such operation, an interpreter has to fetch opcode, its arguments, then it has to dispatch on a jump table (we assume it's an optimized interpreter like Python's) to specific opcode handler, which will then load these two numbers into CPU registers, do addition and store the evaluation result, and then jump back to fetch the next opcode, and then dispatch it, and so on and so forth.
Each individual step like this takes at least 10 cycles (or 25 or 100, depends on complexity) if it's a fancy new big core and can also have hidden latency due to how caches and memory prefetching works. Now, when CPU executes machine code, a modern core can execute multiple additions per cycle, like 4 or even 6 on the newest cores. This alone, means 20-60 times of difference (because IPC is never perfect), and this is with the simplest operation that has just two operands and no data dependencies or other complex conditions.
Once you know this, it becomes easier to reason about overhead of interpreted languages.
I love C sharp and I'm really productive in it, but I've worked at so many places which have tried to get performance out of Python.
> The bodies of many, many PhD students have been hurled unto a big mound to optimize JavaScript within an inch of it's life.
One popular implementation is uvloop (https://github.com/MagicStack/uvloop) which basically just implements the loop using libuv which takes care of doing stuff like `select` as you describe.
Better off, if making some toy loop, to not bother with the forever case.
Adding other resources at the end which explain how it _really_ works under the hood would be great.
See here for a much more in-depth description of coroutines in python,
https://aosabook.org/en/500L/a-web-crawler-with-asyncio-coro...
[1] https://docs.python.org/3/library/typing.html#typing.Generat...
As a caveat, I point out the type signature for academic interest, I prefer to use the simpler Iterable and Awaitable types in type signatures.
Now granted most uses of generators are simpler than that, but still that's a lot of special syntax. And if you are writing Python without type-checking, that's a lot of bookkeeping you'd have to do when reading code.
By the whole thing do you mean the async functionality in Python in general? What is the tragedy?
It feels like the worst of both worlds at times. I think Rust suffers from this too, whereas Node.js has managed to make async the default mode of operation for enough things that it doesn't feel so bad. Especially with Typescript.
https://www.youtube.com/watch?v=bzkRVzciAZg
"Node.js is bad ass rock star tech".
Javascript is inherently async
The tragedy is the whole Python 3 debacle, and if they were going to make a break like that, they should have introduced BEAM-like concurrency instead of async.
For reasons I don't understand, the Python community seems to have an intense phobia and hatred towards threads. I've always used threads instead of async though, and by writing in an Erlang-like style (message passing and no shared mutable data) you can keep things reliable. This has worked ok for me with 1000s of threads, though not millions or anything like that.
Python by now is mostly legacy cruft, it feels to me sometimes. Everything else of course has its own problems.
AFAIK it's not outlined in a PEP, but my intuition is that Python was designed to be accessible and "simple". It wasn't built to be a systems programming language. So for all its faults with the GIL I don't think it should be a major point of contention. For any task it's always about choosing the right tool for the job.
There is considerable effort at the moment to remove the GIL in future versions and implement better threading support. At 3.11/3.12, it has come a long way in terms of performance.
For a simple language, Python has so many different ways to do concurrency, some of which are misleading to new users. Event loop ("async") would've been nice and simple from the start, but they were late to that party. Everything vaguely related to packaging is also confusing, from the __init__.py files to the PIP/Conda/whatever stuff. They had a whole breaking Py2->3 transition and still didn't fix that.
Most concurrent Python code that I've written has not been cpu intensive, FWIW. Just an application talking to a bunch of i/o ports or sockets, but mostly idle. My CPU intensive stuff has tended to be single threaded so I can just launch instances in separate processes. YMMV.
Yeah the tooling situation is crazy enough that I don't begin to understand it.